diff --git a/.gitattributes b/.gitattributes index a6344aac8c09253b3b630fb776ae94478aa0275b..1d79204793eaa0ade9b8cc7b75ac879d6ae05e31 100644 --- a/.gitattributes +++ b/.gitattributes @@ -33,3 +33,27 @@ saved_model/**/* filter=lfs diff=lfs merge=lfs -text *.zip filter=lfs diff=lfs merge=lfs -text *.zst filter=lfs diff=lfs merge=lfs -text *tfevents* filter=lfs diff=lfs merge=lfs -text +image/blobs/sha256/069395655fd57335d9426528f7f49731be6cf7ed963ca649e4ff7c9697d37e0a filter=lfs diff=lfs merge=lfs -text +image/blobs/sha256/0936b6c31255cfe6d029dab00fc251d8771f5e3d228ac7da23e395e1161c6992 filter=lfs diff=lfs merge=lfs -text +image/blobs/sha256/3ac9de4540f40bdf99f054641b2fb36a38f49253d1727fb08a73272d4208ce1b filter=lfs diff=lfs merge=lfs -text +image/blobs/sha256/46f21edbef18d35cb9e2f400eeb6d8f8fac5873f8620e8d96829a9c01447ed52 filter=lfs diff=lfs merge=lfs -text +image/blobs/sha256/98c4455a98982b35380ec62f90d5cf0edff02eeb4b3bb18595d871fdeb0c0c67 filter=lfs diff=lfs merge=lfs -text +image/blobs/sha256/a15921f6aeba449853d8dfc7215c466ad00cff30329183f8b839522d33bc6552 filter=lfs diff=lfs merge=lfs -text +image/blobs/sha256/be04a2af1cfe08908467229c0dd35fc629cb65ba98a497e7b627499f0c76821b filter=lfs diff=lfs merge=lfs -text +image/blobs/sha256/c006f8f8bd1ac6c419a231b5d94521d87d03b141b76873d7a4d7fb68a933e57a filter=lfs diff=lfs merge=lfs -text +image/blobs/sha256/c74a07e1fb0aaf18cdccf3f41385a21dd137109c155363830e2ac2c565d7ac43 filter=lfs diff=lfs merge=lfs -text +image/blobs/sha256/ca0cef561ac457e6f5fe6e064c167bfc523b6e975437121969a8f39630a0a424 filter=lfs diff=lfs merge=lfs -text +image/blobs/sha256/d447e3306f087f0a9b594fb8a92f46880560372c42349c9fc8c10e5bb634aadd filter=lfs diff=lfs merge=lfs -text +image/blobs/sha256/e1ff664083a6df132a0751cb98c17c425fd5908c3434270f38e2de56763f06ba filter=lfs diff=lfs merge=lfs -text +image/blobs/sha256/e324ba29d9a1eb6f52840139e2d929d488845bcadb738439b521f7900e1e22a3 filter=lfs diff=lfs merge=lfs -text +image/blobs/sha256/fc48d390b577ea739097ee841cf6e598acbbe40f0f42571c6cb450d57d7a09f2 filter=lfs diff=lfs merge=lfs -text +media/dp_kashiwanoha_dense_tt_vs_cpu.png filter=lfs diff=lfs merge=lfs -text +media/dp_nuscenes_scene-0061_bev_tt_NC.gif filter=lfs diff=lfs merge=lfs -text +media/dp_nuscenes_scene-0061_kf06_bev_tt_NC.png filter=lfs diff=lfs merge=lfs -text +media/dp_nuscenes_scene-0061_kf18_bev_tt_NC.png filter=lfs diff=lfs merge=lfs -text +media/dp_nuscenes_scene-0061_kf18_cam_front_tt_NC.jpg filter=lfs diff=lfs merge=lfs -text +media/dp_nuscenes_scene-0103_kf12_bev_tt_NC.png filter=lfs diff=lfs merge=lfs -text +media/dp_nuscenes_scene-0103_kf12_cam_front_tt_NC.jpg filter=lfs diff=lfs merge=lfs -text +media/dp_nuscenes_scene-0757_kf11_bev_tt_NC.png filter=lfs diff=lfs merge=lfs -text +media/dp_nuscenes_scene-0916_kf11_bev_tt_NC.png filter=lfs diff=lfs merge=lfs -text +media/dp_straight_road_tt_vs_cpu.png filter=lfs diff=lfs merge=lfs -text diff --git a/OPT_BASELINE.md b/OPT_BASELINE.md new file mode 100644 index 0000000000000000000000000000000000000000..48fc84afa29ae1ebde19e6d4228656d71750df34 --- /dev/null +++ b/OPT_BASELINE.md @@ -0,0 +1,535 @@ +# diffusion-planner-p150 baseline on the p150 (before optimization) + +Date 2026-10-08. Code is the baseline commit `5541833` (first correct port, no optimization): the port's last commit +`34b6049` plus ttaw re-vendored to 0.17.1 and two measurement scripts, `code/scripts/bench.py` (rewritten: the full +stage table) and `code/scripts/profile_ops.py` (new). `git diff 34b6049 5541833` touches no model code (`tt/`, +`host/`, `reference/`, `api.py`, `device.py`, `io.py`, `server/` are unchanged). +tt-metal `44d66500520` (v0.80.0-dev20261006-78) + the ETH-dispatch patch (workspace `patches/tt-metal-eth-dispatch.patch`, +sha256 `08d0ddf6…45cc`, 4 files; `eth_patch: true` in every log). ttaw 0.17.1 @ common `290f22b` +(`code/tt_diffusion_planner/ttaw/VENDORED.json`; `vendor.py --check`: 0 differences). Weights +`AutowareFoundation/diffusion_planner` @ `423efde67f5` (tag `v5.0`). Input: `code/tt_diffusion_planner/samples/kashiwanoha_dense.npz` +(one planning instant on the kashiwanoha test map: 88 neighbours, 123 lanes, 17 route lanes, 60 line strings, ego at +6 m/s), batch 1. The stage bench also runs `straight_road.npz` (shipped) and one public-data instant, nuScenes +v1.0-mini `scene-0103_kf14` (42 neighbours, 103 lanes, 7 route lanes, 10 polygons, 26 line strings; CC BY-NC-SA 4.0, +used locally, never shipped). + +Configuration of every number unless a row says otherwise: ETH dispatch, 1 CQ, 12×10 grid +(`device.compute_with_storage_grid_size()` printed in each log), one p150b (KMD 2.10.0, firmware bundle 19.13.1.0). +AICLK 1350 MHz: sysfs `tt_aiclk` was sampled every 50 ms during every bench, median 1350 MHz and min 1343 MHz. Board +power had a median of 76 W and a max of 86 W, at 63-68 °C. The profiler CSV header also gives 1350 MHz. The host is +shared (AMD EPYC-Rome, 8 vCPUs, up to 8 other agents, load average 8-12 during these runs), so host-side stages are +quoted as p50 / p99 (min). + +## Summary + +| | baseline | +|---|---| +| device latency per plan (one trace replay + sync) | **102.13 ms** p50, 104.81 p99, 102.05 min | +| back-to-back replays (device time per plan) | **102.04 ms** = 9.80 plans/s; the same on all three scenes (dense capacities) | +| `model(inputs=arrays)` end to end | **117.90 ms** p50, 134.56 p99, 112.87 min (`.npz` path: 124.76 ms) | +| programs per plan (device profiler, traced replay) | **6,282** (309 unique programs; the fake-ttnn count was 6,408) | +| device kernel sum / op-to-op gaps / span (profiled replay) | 99.27 / 4.39 / 103.66 ms. Kernel-bound, not gap-bound: the median program runs 5.76 µs, and 4,126 programs under 10 µs add up to 19.2 ms | +| largest single sink | the encoder's channel-MLP and pre-projection matmuls on 4-D `[1, E, T, C]` activations, which run on **4-8 cores**: 90 programs, **25.8 ms** (the 10 slowest programs of the plan are all neighbour-mixer channel matmuls, ~970 µs each) | +| cost of the numerics defaults (PORT_LOG decision 11), re-measured | **+33.4 ms** over the round-1 defaults (68.64 ms b2b) and **+58.3 ms** over the fastest graph (43.74 ms, which fails the gates). Per knob: decoder split matmuls 20.2, fp32 LayerNorm 18.1, fp32 matmul attention 15.5, encoder split matmuls 6.0 ms | +| dispatch / CQ (D14) | ETH-1CQ stays. WORKER is equal (−0.17 ms on 11×10). 2 CQs cost +9.8 ms on ETH (±0 on WORKER) | +| accuracy | 44 device tests, every gate value and all 99 per-scene e2e numbers identical to the port's final run | + +## How to run + +```bash +ROOT=/home/ubuntu/experiments/tt-models; cd $ROOT/bundles/diffusion-planner-p150; source $ROOT/bin/tt-env.sh +export PYTHONPATH=$PWD/code:$PYTHONPATH HF_HUB_OFFLINE=1 # weights: $DIFFUSION_PLANNER_WEIGHTS_DIR > workspace assets > HF cache +# accuracy gates (per-module PCC + end-to-end agreement vs the fp32 CPU reference; 99 scenes with the workspace goldens) +$ROOT/bin/devrun -t 1800 -- env TTAW_GATES_READONLY=1 python -m pytest -q -s \ + code/tt_diffusion_planner/tests/test_pcc_device.py code/tt_diffusion_planner/tests/test_e2e_device.py +# stage breakdown (load / host_pre / pack / host_in / H2D / trace / D2H / host_post / e2e / b2b: p50, p99, min; AICLK) +$ROOT/bin/devrun -t 900 -- python code/scripts/bench.py --iters 100 --json /tmp/bench.json \ + --input code/tt_diffusion_planner/samples/kashiwanoha_dense.npz --input code/tt_diffusion_planner/samples/straight_road.npz \ + --input $ROOT/research/diffusion-planner/public_data/inputs/nuscenes/scene-0103_kf14.npz +# dispatch / CQ matrix: the same per configuration, one process each: --dispatch eth|worker --num-cqs 1|2 +# numerics cost: the same with DIFFUSION_PLANNER_SPLIT_MATMUL / _LN_FP32 / _ATTN_MATMUL set (timing only, see below) +# device profile: one eager plan (stage + layer-kind signposts) and one traced replay, between signposts +$ROOT/bin/devrun -t 3600 -- python -m tracy -r -p -v --op-support-count 16000 --no-web-server \ + -o $ROOT/generated/profiler/diffusion-planner_baseline code/scripts/profile_ops.py +tt-perf-report --start-signpost trace --end-signpost trace_end --arch p150 +python $ROOT/logs/diffusion-planner/baseline/scripts/analyze_profile.py out.json --md out.md +``` + +The job scripts of this baseline are in `logs/diffusion-planner/baseline/scripts/` (workspace): + +- `windowA.sh` (one devrun window, 04:35-04:58 UTC): job 1, the device suite and the alloc-tracking run; then job 2, + the bench and the dispatch / CQ matrix. +- `windowB.sh` (05:54-06:15 UTC): job 3, the numerics ablation; job 4, the Tracy profile; job 5, the + ttnn-visualizer captures. +- `host_suite.sh`, `analyze_profile.py`, `profile_breakdown.py`, `tables.py`, `fake_flops.py`, `visualizer_*.py`. + +The plan issues 6,282 programs, more than the profiler's default 1,000-program buffer. So the profile runs with +`--op-support-count 16000`: 48 B per program per RISC, about 77 MB of DRAM per channel. `ttnn.ReadDeviceProfiler` is +called before and after each section, and nothing was dropped (no full-buffer warning). + +## Accuracy (baseline) + +Re-run on this commit after re-vendoring ttaw 0.11.0 -> 0.17.1, with the gates read-only: + +- device suite: **44 passed** in 48.6 s, and 44 passed again under `TT_METAL_TRACE_ALLOC_TRACKING=1`; +- every gate value and every per-scene number of the 99 scenes is identical to the port's final run (job 19, ttaw + 0.11.0); +- host suite on the fake ttnn: 105 passed, 46 skipped (the 44 device tests, the ORT module, one dead test). + +Logs: `logs/diffusion-planner/baseline/device_suite.log`, `alloc_tracking.log`, `e2e_device.json`, `host_suite.log`. + +| check | gate | baseline | +|---|---|---| +| `enc.ego` / `enc.neighbor` / `enc.lane` / `enc.route` / `enc.polygon` / `enc.line_string` PCC vs the fp32 reference (valid rows, 2 samples) | ≥ 0.999 | 0.999972 / 0.999952 / 0.999992 / 0.999998 / 0.999994 / 0.999998 | +| `enc.goal` / `enc.ego_shape` / `enc.turn` PCC | ≥ 0.999 | 0.999991 / 1.000000 / 0.999995 | +| `enc.encoding` PCC (valid tokens) | ≥ 0.999 | 0.999978 | +| `dec.eval` PCC (teacher-forced decoder, min over the 11 evaluations x 2 samples) | ≥ 0.999 | ≥ 0.9999995 (prints 1) | +| end-to-end ego trajectory vs the reference, max over 99 scenes (2 samples, 5 research scenes, 92 nuScenes-mini instants): max / mean displacement | ≤ 1.0 m / ≤ 0.3 m | 0.313 m / 0.143 m (both on nuScenes scene-0103_kf14) | +| turn-indicator command identical | = 1.0 | 99 / 99 | +| neighbours: median per-agent max displacement, max over scenes | ≤ 1.5 m | 0.086 m | +| CPU reference vs ONNX Runtime on the shipped ONNX (checks the reference itself; research venv, PORT_LOG 5.1) | ≥ 0.9999 | final x0 PCC ≥ 0.9999999999982 (7 scenes), ≥ 0.99999999983 (92 nuScenes instants); 48 tests passed | + +## Performance (warm, batch 1, 100 iterations per stage) + +`code/scripts/bench.py --iters 100` (job 2, `matrix_eth-1cq.{log,json}`): + +- `trace` is one replay of the whole plan plus a device sync, as served; +- `b2b` is the device time per plan over 50 back-to-back replays (3 rounds); +- `e2e` is `model(inputs=)`, and `e2e_path` is `model(inputs=<.npz path>)`. + +Values are ms, p50 / p99 (min). + +| stage | kashiwanoha_dense (shipped default) | straight_road (shipped) | nuScenes scene-0103_kf14 (public data) | +|---|---:|---:|---:| +| `load`: `.npz` decode + schema check (path inputs only) | 7.02 / 18.20 (5.79) | 6.75 / 16.81 (5.86) | | +| `host_pre`: the node's pre-processing | 4.76 / 12.55 (3.84) | 4.32 / 12.58 (3.57) | 4.25 / 6.16 (3.74) | +| `pack`: the 19 persistent trace inputs | 0.45 / 4.52 (0.36) | 0.44 / 1.06 (0.38) | 0.43 / 0.69 (0.34) | +| `host_in`: TILE host tensors | 2.41 / 6.80 (2.02) | 2.41 / 6.79 (2.00) | 2.31 / 3.35 (1.92) | +| H2D (19 tensors, 3.6 MB) + sync | 1.14 / 6.32 (0.90) | 1.09 / 5.43 (0.91) | 1.03 / 2.72 (0.90) | +| **device trace, one blocking plan** | **102.13 / 104.81 (102.05)** | 102.14 / 104.32 (102.07) | 102.13 / 102.19 (102.05) | +| **back-to-back trace replays, per plan** | **102.04** (rounds 102.09, 102.04, 102.04) = 9.80 plans/s | 102.04 | 102.04 | +| D2H (one packed read, 460 KB) | 0.58 / 4.24 (0.31) | 0.59 / 1.50 (0.36) | 0.56 / 0.75 (0.30) | +| `host_post`: the node's post-processing | 4.32 / 11.54 (3.57) | 3.91 / 9.47 (3.29) | 3.93 / 6.96 (3.50) | +| **e2e `model(inputs=arrays)`** | **117.90 / 134.56 (112.87)** | 126.58 / 144.70 (114.23) | 114.52 / 122.91 (113.05) | +| e2e `model(inputs=path)` | 124.76 / 147.74 (118.99) | 121.92 / 143.29 (118.41) | | + +- **Timing breakdown.** `model(...).timing_ms` p50 (kashiwanoha): + - preprocess 4.77 ms; + - device 106.83 ms (`host_in` + H2D + replay + D2H + unpack); + - postprocess 4.36 ms; + - total 117.35 ms. +- **Start-up.** The first call after `from_pretrained` takes 118.2 ms (1.0x). `from_pretrained` takes 11.6 s with a + warm kernel cache: ~5.1 s to open the device, 0.75 s to build (read the ONNX weights and upload them), 5.8 s to warm + up and capture. +- **The device time does not depend on the scene.** All three scenes take 102.13-102.14 ms. Every plan computes the + full capacities whatever the scene holds (no compaction): + - 320 neighbours, 140 lanes, 25 route lanes, 10 polygons and 60 line strings in the encoder; + - 352 decoder rows (321 agents); + - 576 tokens (564 real). +- **Host overhead.** Around the trace the host adds ~15.8 ms per in-process request: pre 4.8 + pack 0.5 + + host_in 2.4 + H2D 1.1 + D2H 0.6 + post 4.3 ms, plus the schema check. Handing over an `.npz` path adds 7 ms. +- **Noise.** The p99 / mean spread of every host stage (and of `e2e`) comes from the shared host. Straight_road's + e2e p50 is higher than kashiwanoha's only because other agents loaded the CPU during that loop. The device numbers + (`trace` min, `b2b`) are stable to ±0.05 ms. +- **Memory.** 309 unique programs; the trace buffers hold 74.6 MB of DRAM; 48.1 MB of weights and constants are + uploaded. + +## Dispatch / CQ matrix (D14) + +One process per configuration, all in one lock window (job 2, `matrix_.{log,json}`); default sample, 100 +iterations. + +| config | grid | trace p50 / p99 (min) | b2b per plan | e2e p50 / p99 (min) | H2D p50 | D2H p50 | load (JIT cache) | first call | +|---|---|---:|---:|---:|---:|---:|---:|---:| +| **ETH-1CQ (default)** | 12x10 | **102.13** / 104.81 (102.05) | **102.04** | **117.90** / 134.56 (112.87) | 1.14 | 0.58 | 11.6 s (warm) | 118.2 | +| ETH-2CQ | 12x10 | 112.02 / 112.96 (111.94) | 111.87 | 125.64 / 141.03 (123.65) | 1.02 | 0.50 | 11.0 s (warm) | 125.0 | +| WORKER-2CQ | 11x10 | 102.03 / 106.28 (101.90) | 101.88 | 126.63 / 142.67 (114.26) | 0.94 | 0.52 | 417.6 s (cold) | 114.7 | +| WORKER-1CQ | 11x10 | 101.97 / 104.53 (101.89) | 101.87 | 116.10 / 135.84 (112.62) | 0.98 | 0.58 | 446.1 s (cold) | 113.6 | +| ETH-1CQ, repeated last | 12x10 | 102.14 / 102.37 (102.07) | 102.04 | 117.35 / 124.73 (114.13) | 1.14 | 0.63 | 9.4 s (warm) | 117.8 | + +- **WORKER vs ETH: equal, so ETH stays the default.** WORKER dispatch is 0.17 ms (0.17 %) faster per plan, on 110 + cores instead of 120. Two effects cancel: + - The 12th column hardly matters to this graph (Grid usage below). The 3,775 programs that do use all 120 cores + are short element-wise programs, bound by a fixed cost per program, not by the core count. The 25.8 ms of 4-8-core + matmuls do not use the grid at all. + - WORKER saves ~0.5 µs of dispatch per program (probe P1). + + That is not "clearly better", so ETH stays the default (D14, `serve.env`). Re-measure once the op count drops, + when the 12th column should start to pay (DISPATCH.md). +- **CQ count: 1.** + - On ETH, a second CQ makes the replay itself +9.8 ms (+9.6 %) per plan, about +1.6 µs per program. + - On WORKER it costs nothing (101.88 vs 101.87 ms). So the penalty belongs to the patched ETH 2-CQ dispatch + topology. It was not profiled here; `--profile-dispatch-cores` would show it. + - What 2 CQs could hide is the upload of the next plan while the trace runs: ~1.1 ms of H2D per 117 ms request, + and only when requests are pipelined. + + Decision: **1 CQ** (the D14 default); `DIFFUSION_PLANNER_NUM_CQS=1` stays pinned. +- A WORKER open recompiles every kernel: 7 min here, for both CQ counts. The served configuration never pays this. +- No drift over the 21-minute window: ETH-1CQ measured 102.13 ms first and 102.14 ms last (trace p50), and 102.04 / + 102.04 ms (b2b). + +## What the numerics defaults cost (precision ablation) + +Job 3 (`precision_.{log,json}`) ran one process per configuration on ETH-1CQ with the default sample: 30 +iterations, and b2b over 3 x 30 replays. The configurations differ only by the `DIFFUSION_PLANNER_*` knobs. **This +is timing only: no gate was run here, and only `shipped` is gate-clean.** The PORT_LOG evidence for three of the +others: + +- the round-1 defaults fail `ego.mean_err_m` on nuScenes scene-0103_kf14 (0.347 m > 0.3, job 11); +- bf16 SDPA instead of the fp32 matmul attention (PORT_LOG's "decoder split alone", the same knobs) gives 0.74 / + 0.33 m (max / mean) on scene-0103_kf14, over the 0.3 m mean gate (job 13, `logs/diffusion-planner/split_error_decsplit.json`); +- the fastest graph (fused LN, no split, SDPA) fails `enc.ego` (0.99856 / 0.99261, job 4). + +The decoder-split, encoder-split and fused-LN removals were never gated. + +| config | `SPLIT_MATMUL` | `LN_FP32` | `ATTN_MATMUL` | trace p50 | b2b | vs shipped | unique programs | trace buffers | +|---|---|---|---|---:|---:|---:|---:|---:| +| **shipped** (decision 11) | `enc.island.*,enc.pre.*,dec.*` | `enc.mixer.*,dec.*` | `enc.fusion.attn,dec.*` | 102.11 | 102.04 | | 309 | 74.6 MB | +| bf16 SDPA instead of fp32 matmul attention | same | same | none | 86.64 | 86.57 | −15.47 | 305 | 69.9 MB | +| no split matmuls in the decoder | `enc.island.*,enc.pre.*` | same | same | 81.93 | 81.85 | −20.19 | 283 | 50.4 MB | +| no split matmuls in the encoder | `dec.*` | same | same | 96.08 | 96.02 | −6.02 | 209 | 72.9 MB | +| fused `ttnn.layer_norm` everywhere | same | none | same | 84.03 | 83.94 | −18.10 | 304 | 43.9 MB | +| round-1 defaults (PORT_LOG job 11: 68.7 ms) | `enc.island.*,enc.pre.*,dec.preproj.*` | `enc.mixer.*,dec.*` | none | 68.72 | 68.64 | **−33.40** | 287 | 50.4 MB | +| fastest graph (PORT_LOG job 4: 43.7 ms) | none | none | none | 43.79 | 43.74 | −58.30 | 173 | 16.2 MB | + +- **Precision cost of the shipped defaults.** It is **33.4 ms** over round 1, which reproduces PORT_LOG's 102.1 vs + 68.7 ms. Over the fastest graph it is 58.3 ms. +- **The knobs add up almost exactly.** The four single-knob removals sum to 59.8 ms, against 58.3 ms with all of + them off. +- **On the card this cost is the first optimization target:** it is the price of the accuracy the nuScenes e2e gate + needs. Items 2-4 below recover it with fused kernels that keep the fp32-level accuracy. They do not relax precision: + precision policy §9.3, and every numerics change re-checked on all 99 scenes. Item 1 is a separate sink of similar + size that this profile found, unrelated to precision. + +## Host vs device, and every host<->device transfer per plan + +Host, per plan: numpy, the node's own pre- and post-processing (the PLAN.md 2.12 host rows). There is no host +fallback inside the plan. + +- `load`: decodes the `.npz` and checks the 15 raw tensors against `INPUT_SCHEMA` (names, shapes, finite values). + Only `model(inputs=)` does this; the server decodes its JSON / base64 body instead. +- `host_pre` (`host.prepare`): + - normalization (all-zero rows kept) and the speed masks; + - the encoder's pre-matmul input plumbing: history truncation, velocity zeroing, validity masks, position features + with the ONNX `atan2` decomposition and the pseudo-heading quirk, point deltas; + - the decoder masks and `x_T`. +- `pack` (`tt.inputs.plan_inputs`) builds the 19 persistent trace inputs: the input columns of the exact affine + rewrites, the 576-token arrays, the key-bias rows, and `cs` / `y0` on 352 rows. +- `host_in`: converts them to ttnn host tensors (TILE layout, host tilization). +- `host_post` (`host.make_output`): denormalization, trajectory velocity / force-stop / acceleration, the predicted + paths of the non-empty neighbour rows, the turn-indicator decision. + +Device: everything else, as ONE metal trace (`plan`): + +- the encoder: 6 mixer trunks, entity heads, small encoders, token assembly, 6 fusion blocks; +- the hoisted cross K / V; +- 11 x (DiT evaluation + fp32 DPM-Solver++(2M) update + prefix constraint); +- the turn head and the output pack. + +**H2D per plan: 19 `copy_host_to_device_tensor` into the persistent inputs** (CQ0; with 2 CQs, CQ1 + events): 1.19 MB +logical, 3.60 MB as TILE-padded fp32 / bf16. + +| input | shape | dtype | logical B | TILE-padded B | +|---|---|---|---:|---:| +| `neighbor_x` | [1, 320, 6, 9] | fp32 | 69,120 | 1,310,720 | +| `lane_x` | [1, 140, 20, 8] | fp32 | 89,600 | 573,440 | +| `cs` (current states in the t = 0 columns) | [1, 1, 352, 324] | fp32 | 456,192 | 495,616 | +| `y0` (`x_T * mask0`; zeros at temperature 0) | [1, 1, 352, 324] | fp32 | 456,192 | 495,616 | +| `line_string_x` | [1, 60, 20, 6] | fp32 | 28,800 | 245,760 | +| `route_x` | [1, 25, 20, 8] | fp32 | 16,000 | 102,400 | +| `polygon_x` | [1, 10, 40, 5] | fp32 | 8,000 | 81,920 | +| `token_valid` | [1, 1, 576, 1] | fp32 | 2,304 | 73,728 | +| `pos_aug` | [1, 1, 576, 15] | fp32 | 34,560 | 73,728 | +| `neighbor_aux` | [1, 1, 320, 4] | fp32 | 5,120 | 40,960 | +| `fusion_key_row` | [1, 1, 1, 576] | bf16 | 1,152 | 36,864 | +| `agent_key_row` | [1, 1, 1, 352] | bf16 | 704 | 22,528 | +| `lane_aux` | [1, 1, 140, 29] | fp32 | 16,240 | 20,480 | +| `ego_x`, `static_x`, `route_aux`, `goal_x`, `ego_shape_x`, `turn_x` | small | fp32 | 3,344 | 6 x 4,096 | + +**D2H per plan: 1 packed read** (`pack_outputs`), 117,664 fp32 = 470,656 B in one row: + +- `final_x0` [352, 324]; +- the turn logits [5]; +- the ego rows of the 11 published iterates [11, 324]. + +There is no mid-graph host round trip and no host fallback op. + +## Grid usage + +`tt/` sets no core grid or program config on any op. The only grid consumer, C20's SDPA config, reads +`device.compute_with_storage_grid_size()`, and the default fp32 matmul attention does not use it. So every op picks +its own core count: + +| core count per program (traced replay) | programs | kernel ms | what | +|---|---:|---:|---| +| 120 | 3,775 | 39.4 | element-wise / typecast / reduce programs, spread over all cores whatever their size | +| 64-119 | 1,050 | 14.1 | decoder matmuls on 66-72 cores, softmax on 88, token-mixing matmuls on 100-117 | +| 16-63 | 708 | 11.9 | decoder matmuls on 48 cores (561 programs, 9.7 ms) | +| **2-15** | **712** | **33.6** | **matmuls on 4 / 8 cores: 148 programs, 26.0 ms, of which the encoder's batched `[1, E, T, C] @ [C, 128]` channel matmuls are 90 programs, 25.8 ms**; the attention P·V matmuls on 11 cores (66 programs, 4.6 ms); decoder LN reductions on 11 cores | +| 1 | 37 | 0.2 | slices, small reshapes | + +Matmul efficiency is low everywhere. tt-perf-report gives a weighted mean of 7.2 % of the FLOP roofline over the +1,365 matmuls (max 12.9 %), and 6.8 % of the DRAM roofline (35 GB/s) over the modelled ops. + +## What is already fused / traced + +- **The whole plan is one trace:** encoder + 11 DiT evaluations + 10 solver updates + turn head + pack, replayed + with persistent inputs. Replay == eager bit for bit (`test_plan_replay_equals_eager`, and the profile run). The + device suite is green under `TT_METAL_TRACE_ALLOC_TRACKING=1`. +- **Exact rewrites already in the graph** (`tt/params.py`, proven on the host): + - the cross-attention K / V of the 3 DiT blocks is hoisted out of the solver loop (once per plan, not 11 times); + - the per-step adaLN is folded into the LayerNorm affine rows (no adaLN MLP on the device); + - the solver update is folded into `y' = A y - B m_k + Cm m_(k-1)`, with the prefix constraint as `x = y + cs` and + the t = 0 output columns zeroed in the last projection's weights; + - the ego / neighbour pre-projection runs as a pad-relative fp32 island; + - the small embeddings (neighbour type, lane speed / attributes, position) are one matmul over host-built columns; + - the turn head is `W_sel` / `W_pool / 564`, and the agent embedding is folded into bias rows. +- **Fused in stock ops:** + - bias + GELU / tanh-GELU in `ttnn.linear` (2-D inputs); + - Q | K | V in one linear and one `nlp_create_qkv_heads` (DiT self-attention); + - K | V in one linear + `split_q_kv` (fusion); + - one packed readback. +- **Not fused: the numerics defaults** (PORT_LOG decision 11), counted in programs per call: + - each split hi / lo matmul is **8.3 programs** on average instead of 1: typecast, typecast, subtract, 3 + matmuls / linears, 2 adds, then GELU / typecast as needed; + - each fp32 LayerNorm is **9 programs** instead of 1: mean, subtract, multiply, mean, add, rsqrt, multiply, then + multiply / add for the affine; + - each fp32 attention is **5 programs** (matmul, scale, mask add, softmax, matmul), plus the head split / merge. +- **Not fused: glue.** Residual adds, adaLN gate multiplies, the mixer token-mixing transposes, the solver update + (3-5 element-wise programs per step) and the ego-row slices. +- **Not done at all:** + - every activation is DRAM-interleaved: no sharding, no L1 residency; + - no custom kernel, no `generic_op`, no megakernel; + - no compaction: all 321 agent rows and 564 tokens are computed whatever the scene holds. + +## Device profile (traced replay) + +The profile (job 4, `profile_ops.py`) has two sections: + +- one eager plan with stage (`m:`) and layer-kind (`c:`) signposts: 778.6 ms, dispatch-bound at a median of 101 µs + per eager op; +- one traced replay of the same 6,282 programs. The op codes match the eager run one by one, so every replayed program + inherits its stage and kind. + +The replay with the profiler on took 105.13 ms on the host, against 102.1 ms without it. The eager run after the +capture logs tt-metal's "allocating device buffers ... active trace" warning once; `run_eager` frees its buffers +before the replay, and the replay equals the eager plan bit for bit. + +| | traced replay | +|---|---| +| programs per plan | **6,282** | +| device kernel sum | **99.27 ms** | +| op-to-op gaps (sum / median / p90 / max) | **4.39 ms** / 0.56 µs / 0.65 µs / 7.97 µs | +| span (first FW start to last FW end, 1350 MHz) | **103.66 ms** | +| kernel time per program, p10 / p50 / p90 / p99 / max | 2.7 / 5.76 / 32.2 / 96.6 / 984 µs | +| programs under 5 µs / under 10 µs | 2,673 / 4,126 (19.2 ms together) | +| math fidelity | HiFi4 on all 6,106 compute programs | + +The plan is **kernel-bound**: gaps are 4 % of the span. The time goes to a few slow, badly parallelised programs and +to thousands of small ones, each costing ~5-8 µs whatever its size. The FLOPs are minor: the device graph does 165 +GFLOP of matmul per plan with every split pass counted (132 GFLOP of it in the decoder; `fake_flops.log`). That is +~1 ms at the 120-core HiFi4 peak (~162 TFLOPS, TT_PLATFORM.md). + +**By stage** (kernel + gaps, ms): + +| stage | programs | ms | share | +|---|---:|---:|---:| +| encoder: 6 mixer trunks (pre-projection 8.8, 6 x 6 MixerBlocks 39.2, pool + heads 0.3) | 1,274 | 48.35 | 46.6 % | +| encoder: fusion (6 blocks + final LN), small encoders, tokens, masks | 149 | 3.23 | 3.1 % | +| decoder: 11 evaluations, 4.67 ms each incl. its solver update (433-443 programs each; e00-e09 4.38-4.40 ms kernel, e10 4.45) | 4,791 | 51.42 | 49.6 % | +| cross K / V hoist, turn head, output pack | 68 | 0.66 | 0.6 % | + +**By layer kind and part** (kernel + gaps, ms; `kind_by_group.md`): + +| part | split matmul | plain linear | fp32 LN | fp32 attention | glue | heads / mask / fused LN | total | +|---|---:|---:|---:|---:|---:|---:|---:| +| encoder mixer trunks | 8.31 (24 calls, 346 µs each) | **23.27** (159 calls) | 12.96 (72 calls, 180 µs each) | | 3.77 | 0.04 | 48.35 | +| encoder fusion + small | | 0.72 | | 1.96 (6 calls) | 0.21 | 0.33 | 3.23 | +| decoder x11 + solver | **26.09** (308 calls, 85 µs each) | | 8.59 (165 calls, 52 µs each) | **14.47** (66 calls, 219 µs each) | 1.34 | 0.93 | 51.42 | +| cross K / V, turn, pack | 0.38 | 0.01 | | | 0.24 | 0.02 | 0.66 | +| **all** | **34.78** | **24.00** | **21.56** | **16.43** | **5.56** | 1.33 | **103.66** | + +**One DiT evaluation** (e5, 4.67 ms): + +| | ms | share | +|---|---:|---:| +| split matmuls (pre-projection, self Q / K / V + out, cross Q + out, 2 MLPs, final) | 2.37 | 50.8 % | +| self- and cross-attention as fp32 matmuls (3 x 0.20 + 3 x 0.24) | 1.32 | 28.2 % | +| fp32 LayerNorms (12 x 48.5 µs + the final layer's 3 LNs 197 µs) | 0.78 | 16.7 % | +| glue, head split / merge, solver update | 0.20 | 4.3 % | + +**By op code:** + +| op | programs | kernel ms | share | +|---|---:|---:|---:| +| Matmul | 1,365 | 52.34 | 52.7 % | +| BinaryNg (element-wise binary) | 2,948 | 31.48 | 31.7 % | +| Reduce (LN means) | 480 | 3.69 | 3.7 % | +| Softmax | 72 | 3.47 | 3.5 % | +| Unary (GELU, rsqrt) | 428 | 2.81 | 2.8 % | +| Typecast | 676 | 2.41 | 2.4 % | +| Transpose (mixer token mixing) | 84 | 1.34 | 1.4 % | +| NlpCreateHeads / NLPConcatHeads | 150 | 1.00 | 1.0 % | +| other (reshape, untilize, LayerNorm, repeat, tilize, concat, slice, pad) | 79 | 0.73 | 0.7 % | + +**Top-10 device programs** (traced replay, single programs). All ten are the neighbour mixer's channel-MLP matmuls +`[1, 320, 64, 128] @ [128, 128]` (6 MixerBlocks x 2 linears), each on **8 cores**: + +| # | op | stage | in0 @ in1 | dtypes in0 / in1 / out | cores | kernel µs | +|---:|---|---|---|---|---:|---:| +| 1 | Matmul (`ch2`) | `enc.neighbor.mix` | [1,320,64,128] @ [1,1,128,128] | bf16 / bf16 / fp32 | 8 | 984.1 | +| 2 | Matmul (`ch2`) | `enc.neighbor.mix` | same | bf16 / bf16 / fp32 | 8 | 977.6 | +| 3 | Matmul (`ch2`) | `enc.neighbor.mix` | same | bf16 / bf16 / fp32 | 8 | 972.8 | +| 4 | Matmul (`ch2`) | `enc.neighbor.mix` | same | bf16 / bf16 / fp32 | 8 | 971.8 | +| 5 | Matmul (`ch2`) | `enc.neighbor.mix` | same | bf16 / bf16 / fp32 | 8 | 971.2 | +| 6 | Matmul (`ch2`) | `enc.neighbor.mix` | same | bf16 / bf16 / fp32 | 8 | 970.1 | +| 7 | Matmul (`ch1`, + GELU) | `enc.neighbor.mix` | [1,320,64,128] @ [1,1,128,128] | fp32 / bf16 / bf16 | 8 | 966.4 | +| 8 | Matmul (`ch1`) | `enc.neighbor.mix` | same | fp32 / bf16 / bf16 | 8 | 965.2 | +| 9 | Matmul (`ch1`) | `enc.neighbor.mix` | same | fp32 / bf16 / bf16 | 8 | 965.2 | +| 10 | Matmul (`ch1`) | `enc.neighbor.mix` | same | fp32 / bf16 / bf16 | 8 | 965.0 | + +**Top op signatures by total time** (`profile_breakdown.md`). Only matmul shapes are shown: element-wise programs are +shared across shapes, and the profiler's op-info cache reports the first call's shapes for them. + +| op | where | kind | in0 @ in1 | cores | programs | ms | mean µs | +|---|---|---|---|---:|---:|---:|---:| +| Matmul | `enc.neighbor.mix` channel MLP | linear | [1,320,64,128] @ [128,128] | 8 | 12 | 11.64 | 970 | +| BinaryNg | mixer blocks | fp32 LN | | 120 | 360 | 9.92 | 27.6 | +| Matmul | decoder MLP fc2 (mlp1 + mlp2, 3 split passes each) | split | [1,1,352,1024] @ [1024,256] | 48 | 198 | 5.71 | 28.8 | +| Matmul | `enc.lane.mix` channel MLP | linear | [1,140,64,128] @ [128,128] | 8 | 12 | 5.09 | 425 | +| BinaryNg | decoder attention scale / mask add | fp32 attention | | 120 | 132 | 4.72 | 35.8 | +| Matmul | decoder MLP fc1 (mlp1 + mlp2, 3 split passes each) | split | [1,1,352,256] @ [256,1024] | 66 | 198 | 3.39 | 17.1 | +| Matmul | `enc.neighbor.pre` (pad-relative island) | split | [1,320,6,128] @ [128,128], [1,320,6,9] @ [9,128] | 4 | 6 | 3.15 | 525 | +| Matmul | decoder self-attention P·V | fp32 attention | [1,8,352,352] @ [1,8,352,32] | 11 | 33 | 2.33 | 70.7 | +| Matmul | decoder cross-attention P·V | fp32 attention | [1,8,352,576] @ [1,8,576,32] | 11 | 33 | 2.26 | 68.4 | +| Matmul | `enc.line_string.mix` channel MLP | linear | [1,60,64,128] @ [128,128] | 8 | 12 | 2.19 | 183 | +| Softmax | decoder cross-attention | fp32 attention | | 88 | 33 | 1.83 | 55.4 | +| BinaryNg | mixer residual adds | glue | | 120 | 72 | 1.75 | 24.3 | +| Reduce | mixer LN means | fp32 LN | | 120 | 72 | 1.39 | 19.3 | + +Compare the 2-D token-mixing matmul of the same trunk: `[1,1,40960,64] @ [64,64]` runs on 117 cores in 69 µs. The +4-D channel MLP with twice the FLOPs takes 970 µs on 8 cores: 14x the time, about 7x slower per FLOP. + +Artifacts: + +- **Profile:** `generated/profiler/diffusion-planner_baseline/reports/2026_10_08_06_13_03/`, holding + `ops_perf_results_*.csv`, the Tracy file and `profile_log_device.csv.zst` (7.3 GB raw; `zstd -d` it before loading + the folder in ttnn-visualizer), plus `.logs/cpp_device_perf_report.csv`. +- **ttnn-visualizer memory / graph reports** (not for timing): `generated/ttnn_visualizer/diffusion-planner_baseline/{decode_once,encoder_taps}/`: + - one decoder evaluation: 508 ops, 20.9 MB `db.sqlite`; + - the encoder: 1,483 ops, 63.5 MB. + - No whole-plan capture was made. At the measured ~1.3k buffer rows per op it would be ~8M rows (a ~2.5 GB capture + JSON) for 11 repeats of the same evaluation. The two captures cover every distinct layer of the plan except the + solver update, the turn head and the output pack. + - Open them with `ttnn-visualizer --profiler-path --performance-path `. + - Capture: `logs/diffusion-planner/baseline/scripts/visualizer_capture.py` (job 5). Import: + `visualizer_import.py`, on the CPU. + +## Ranked optimization opportunities + +Gains are estimated on the 102.0 ms replay and come from the profile above. They are **not additive**: each one +shrinks the base of the next. Every item must keep the frozen gates, re-checked on all 99 e2e scenes (PORT_LOG 8: +the plan's sensitivity to device numerics is chaotic per scene). The precision cost of decision 11 (33.4 ms over round +1) is recovered by items 2-4, not by relaxing precision. + +1. **Run the encoder's channel MLPs on 2-D activations (≈ −20 to −22 ms; exact rewrite, low effort).** + - *What is slow.* The mixer `channels_mlp` (`ch1` / `ch2`) and the channel pre-projections run `ttnn.linear` on + 4-D `[1, E, T, C]` activations. That is a batched matmul on **4-8 cores**: 90 programs, 25.8 ms. + - *MixerBlocks.* `[1, E, 64, 128]` -> `[1, 1, E·64, 128]` is a free view (64 rows = 2 whole tiles per entity). + Then the matmul can spread over the grid, as the token-mixing matmuls already do (69 µs on 117 cores for the + neighbours). + - *Pre-projections.* T = 6 / 20 / 40 rows do not fill a tile, so the view is not free. Two options: + - have `pack` lay the inputs out as `[1, 1, E·T, C]` and re-tile once before the token transpose; + - or give the 4-D matmul an explicit multi-core program config, with the HiFi4 compute config passed too (a + `program_config` alone falls back to LoFi, PLAN 0.2). + - *Estimate.* The channel matmul at 2x the measured token-MLP time of its trunk saves ~17 ms on the MixerBlocks + and ~5 ms on the pre-projections. The math is unchanged (the K = 128 reduction per output). The gates must still + be re-run, because accumulation order may differ. + - *Generalisation.* No `[1, E, ...]`-batched `ttnn.linear` anywhere (the PR:P12 rule extended from token mixing + to channel mixing). +2. **A fused split (bf16x3) matmul, one program per linear (≈ −17 to −19 ms).** + - *Today.* 338 split linears, 2,802 programs, 34.8 ms: decoder 26.1, encoder 8.3, cross K/V 0.4. + - *The kernel.* A `generic_op` matmul that reads the fp32 activation once, splits hi / lo on the fly (SFPU), and + accumulates `x_hi W_hi + x_hi W_lo + x_lo W_hi` in fp32 DEST. + - *What it removes.* All typecast / subtract / add programs and their gaps, 9.3 ms in the decoder. The three + passes also share operand reads: estimated at 1.5x one pass, 8 instead of 16 ms of decoder matmul. + - *Accuracy.* The same ~1e-5 relative error as today's 3-matmul form. + - *Cheaper first step.* Explicit program configs for the 352-row decoder matmuls, which run on 48-72 cores at + ~7 % FLOP utilisation. Pass the compute config with them: a `program_config` without one silently falls back to + LoFi (PLAN 0.2). +3. **A fused fp32 LayerNorm, one program per LN (≈ −15 to −16 ms).** + - *Today.* 237 LNs, 2,133 programs, 21.6 ms: mixers 72 x 180 µs on `[E, 64, 128]`; decoder 165 x 52 µs on + `[352, 256]`. + - *The kernel.* One program per LN (mean / variance in fp32 by Welford or two passes in L1, rsqrt, affine with + the folded adaLN rows), at ~2 element-wise passes: ~45 µs in the mixers and ~12 µs in the decoder. + - *Limit.* The fused `ttnn.layer_norm` would save 18.1 ms (ablation), but it is ~2.5e-3 relative (PR:P10). On the + offset-dominated mixer rows it reaches rel-L2 0.03, and the teacher-forced neighbour mixer drops to PCC 0.9978 + (PORT_LOG job 5). +4. **Attention (≈ −5.5 ms with stock ops, ≈ −12 to −13 ms with a fused kernel).** + - *Today.* 72 fp32 matmul attentions, 16.4 ms, ~228 µs per call. The bf16 SDPA would cost ~13 µs per call + (ablation: −15.5 ms), but fails the e2e mean gate. + - *(a) Stock ops.* Fold the 1/√32 scale into `W_q` / `b_q`: fp64 fold, then the hi / lo split; gate it. Use a + fused scale-mask-softmax for the mask add. Together this removes the 132 decoder (+ 12 fusion) scale / mask + programs: 4.7 + 0.7 ms. + - *Parallelism.* The P·V matmuls (`[1, 8, 352, Sk] @ [1, 8, Sk, 32]`) run on 11 cores (66 programs, 4.6 ms); a + head-parallel program config would spread them. + - *(b) A fused fp32-accurate flash-attention kernel* (`generic_op`): scores stay in L1 / DEST, `Q·Kᵀ` and `P·V` + use hi / lo operands, and the mask comes from the key-bias row. At ~3-4x the bf16 SDPA's cost (~50 µs per call) + it saves ~12.8 ms and supersedes (a). +5. **Glue (≈ −3 to −4 ms).** + - *Mixers: 3.8 ms.* Token-mixing transposes (84 programs, 1.3 ms) and residual adds (72 programs, 1.8 ms). The + entity-parallel mixer kernel of PLAN 5.3 (no transposes) or residual adds in the matmul epilogue remove them. + - *Decoder.* Fold the per-step adaLN gates exactly into per-step copies of the gated weights (`attn.out`, + `mlp1.fc2`): this removes 66 gate multiplies, at ~65 MB of extra DRAM for the 11 x 3 copies with hi / lo parts. + - *Solver.* The update `A y - B m_k + Cm m_(k-1)` becomes one ternary / `generic_op` program instead of 3-5 (83 + programs, 0.45 ms). +6. **Megakernels (D19; after 1-4, gains measured against what is left).** + - **MK-D, the persistent DiT-evaluation kernel.** + - *Today.* One evaluation is 4.67 ms in ~435 programs. Its matmul work is 12.0 GFLOP, split passes included + (74 µs at the HiFi4 peak). + - *Fits in L1.* Its weights are 8.7 MB (DiT blocks) + 1.8 MB (pre-projection, final layer) as bf16 hi parts, plus + the fp32 lo parts; the activations are `[352, 256]` fp32 (360 KB); the hoisted cross K / V is 3.5 MB. + - *Parameters.* Compile-time: agent bucket (352), heads 8 x 32, MLP 1024, depth 3, NFE 11. The folded adaLN + rows and the solver coefficients move from 11 unrolled trace copies into an L1 table (RT-dev constants per + steps value). + - *Target.* 0.5-1 ms per evaluation would cut the decoder from 51 to ~6-11 ms. + - **MK-E (one mixer trunk) and MK-F (one fusion block)** for the encoder. + - **The single persistent megakernel for the whole plan** (PLAN 5.3) is plausible. The 48 MB of uploaded weights / + constants and the largest activation (neighbour mixer `[320, 64, 128]` fp32, 10.5 MB) fit the ~175 MB of + aggregate L1. The compute floor is ~1 ms. + - It ships only if it is faster with the gates intact; the attempt is recorded either way (D19). +7. **Compaction (LOAD-time buckets; 20-40 % of what remains after 1-4, to be measured).** + - The device time is scene-independent: kashiwanoha uses 88 / 320 neighbours, 123 / 140 lanes and 89 / 352 + decoder rows. + - Valid-entity buckets need one trace per bucket and the exact compaction rewrites of PLAN 5.3: + - they shrink the bandwidth-bound encoder work in proportion; + - they shrink the attention scores quadratically (`[8, 352, 352]` -> `[8, 96, 96]` at 89 agents); + - they help the latency-bound decoder matmuls little. +8. **Host side, e2e only (≈ −4 to −6 ms per request).** 15.8 ms of host work and transfers surround the 102 ms replay. + - Pack `neighbor_x` as a 2-D `[1, 1, 1920, 9]` input: 1.31 MB of TILE padding becomes 245 KB, which also feeds + item 1. + - Skip the `y0` upload while `x_T` is unchanged (zeros at temperature 0, the node's default) and send `cs` as its + 4 columns. Together: H2D 3.6 -> ~1.6 MB, and less host tilization (`host_in` 2.4 ms). + - Vectorise the per-entity loops of `host_pre` (4.8 ms) and `host_post` (4.3 ms: the predicted paths of up to 320 + agents). + - The 7 ms `.npz` decode is the caller's choice (pass arrays). +9. **Dispatch configuration: no gain today; re-measure after 1-5.** + - Gaps are only 4.4 ms (0.56 µs median). What pays is removing programs: each removed short program saves its + ~5-8 µs kernel plus ~0.6 µs. + - Re-run the ETH / WORKER and 1 / 2 CQ matrix once the op count drops below ~2,000. ETH's 12th column matters only + once ops become throughput-bound. + - The ETH-2CQ penalty (+1.6 µs per program) needs a dispatch-core profile before 2 CQs are reconsidered. + +## Raw logs and CSVs + +Under `logs/diffusion-planner/baseline/` (workspace) unless absolute: + +| what | files | +|---|---| +| re-vendoring + suites | `host_suite.log` (105 passed / 46 skipped; again on the final tree with the new scripts: `host_suite_final.log`), `device_suite.log` (44 passed), `alloc_tracking.log` (44 passed), `e2e_device.json` (99 scenes), `windowA.log` | +| stage bench + dispatch / CQ matrix | `matrix_{eth-1cq,eth-2cq,worker-2cq,worker-1cq,eth-1cq-repeat}.{log,json}`, `matrix.rcs`, `tables.md` | +| numerics ablation | `precision_{shipped,no_attnmm,no_decsplit,no_encsplit,no_ln32,round1,fastest}.{log,json}`, `precision.rcs`, `windowB.log` | +| device profile | `profile_tracy.log`, `profile_run.json`; ops CSV `/home/ubuntu/experiments/tt-models/generated/profiler/diffusion-planner_baseline/reports/2026_10_08_06_13_03/ops_perf_results_2026_10_08_06_13_03.csv`; summaries `profile_summary.{json,md}` (`analyze_profile.py`), `profile_breakdown.md`, `kind_by_group.md`; tt-perf-report `tt_perf_report_trace.{txt,csv}`, `tt_perf_report_trace_summary.{csv,png}` | +| ttnn-visualizer | `visualizer_capture.log`, `visualizer_import.log`; reports in `/home/ubuntu/experiments/tt-models/generated/ttnn_visualizer/diffusion-planner_baseline/` | +| op counts / FLOPs on the fake ttnn (CPU) | `fake_signposts.log` (6,408 fake ops by stage / kind), `fake_flops.log` (165 GFLOP per plan); the port's per-configuration counts: `logs/diffusion-planner/opcount_fake_r2.log` | +| scripts | `scripts/` (`windowA.sh`, `windowB.sh`, `job{1..5}_*.sh`, `chain.sh`, `analyze_profile.py`, `profile_breakdown.py`, `tables.py`, `fake_*.py`, `visualizer_*.py`) | diff --git a/OPT_REPORT.md b/OPT_REPORT.md new file mode 100644 index 0000000000000000000000000000000000000000..8656ce8a4821b9fe4737f52a16d79080ab34a366 --- /dev/null +++ b/OPT_REPORT.md @@ -0,0 +1,142 @@ +# diffusion-planner-p150 optimization report (p150, ETH dispatch, 12×10 grid) + +**Status: baseline port; optimization pending.** This is the first public release: the functional port, measured +once (`OPT_BASELINE.md`, baseline commit `5541833`, 2026-10-08) and not optimized yet. No optimization round has run, +so there is no step, no rejected attempt and no hang to report, and every number below is the baseline. The whole +plan (encoder, the 11 DiT evaluations with the DPM-Solver++(2M) updates, the turn-indicator head) runs as one metal +trace of 6,282 programs: 102.04 ms per back-to-back replay (9.80 plans/s of device throughput), 117.9 ms per +synchronous `model()` call on the shipped sample. The trace is kernel-bound (op-to-op gaps 4.4 ms of a 103.7 ms span). +Two sinks of similar size come first: the **numerics defaults** that the end-to-end gates need (split hi / lo matmuls, +fp32 LayerNorm, fp32 matmul attention) cost **33.4 ms** over the first device round, and the encoder's channel-MLP and +pre-projection matmuls run on only 4-8 cores (**25.8 ms**, unrelated to precision). + +All numbers: `code/scripts/bench.py` (median of 100 warm iterations, batch 1, +`code/tt_diffusion_planner/samples/kashiwanoha_dense.npz`), ETH dispatch, 1 CQ, 12×10 grid, the pinned numerics +(`DIFFUSION_PLANNER_SPLIT_MATMUL=enc.island.*,enc.pre.*,dec.*`, `_LN_FP32=enc.mixer.*,dec.*`, +`_ATTN_MATMUL=enc.fusion.attn,dec.*`), unless marked otherwise. Device profile: `code/scripts/profile_ops.py` under the +device profiler (one traced replay between signposts). Accuracy gates: `OPT_BASELINE.md` "How to run". + +## Summary + +| | baseline `5541833` (2026-10-08) | **final (= baseline: no round yet)** | +|---|---|---| +| device trace, one blocking plan | 102.13 ms (p99 104.81) | same | +| back-to-back traces | 102.04 ms (9.80 plans/s) | same | +| e2e `model(inputs=arrays)` p50 / p99 | 117.90 / 134.56 ms | same | +| e2e `model(inputs=<.npz path>)` p50 | 124.76 ms | same | +| host pre-processing / pack / host tensors / H2D / D2H / post-processing | 4.76 / 0.45 / 2.41 / 1.14 / 0.58 / 4.32 ms | same | +| device programs per plan (unique programs) | 6,282 (309) | same | +| kernel sum / op-to-op gaps / span (profile) | 99.27 / 4.39 / 103.66 ms | same | +| encoder / 11 decoder evaluations (+ solver updates) | 51.6 / 51.4 ms (4.67 ms per evaluation) | same | +| accuracy gates (PCC / agreement vs the fp32 CPU reference) | 44 / 44 device tests: module PCC ≥ 0.999952, encoding 0.999978, decoder evaluation ≥ 0.9999995; 99 scenes: ego max 0.313 m / mean 0.143 m (gates 1.0 / 0.3 m), turn command 99 / 99, neighbours ≤ 0.086 m (gate 1.5 m) | same | +| served `/predict` `timing_ms.total`, median of 50 (uvicorn on the host, the shipped sample) | 123.1 ms, measured on `c0d84f9` (same device code; the release verification) | same | +| `from_pretrained` load: empty JIT cache / warm cache | 315 s / 8.6 s (`c0d84f9`) | same | + +The host is shared with other agents' jobs, so the host-side rows move with its load (several ms); the device rows +repeat to ±0.05 ms. AICLK 1350 MHz (min 1343) during the stage bench. The release re-check on `c0d84f9` +(`VERIFICATION_2026-10-08.md`) reproduced the device rows: back-to-back replay 102.03 vs 102.04 ms, one plan 102.10 vs 102.13 ms; the host-bound end to end 115.6 vs 117.9 ms p50 moves with the host load. + +## Steps (chronological; each row is one commit) + +| step | commit | what | trace / b2b ms | e2e ms | accuracy | revert switch | +|---|---|---|---|---|---|---| +| baseline | `5541833` | first correct port: the whole plan in one trace, ETH 1CQ 12×10, the numerics defaults of PORT_LOG decision 11, host pre- and post-processing | 102.13 / 102.04 | 117.90 | all gates (above) | – | + +Later commits up to the first release changed no model numerics and no device code: the baseline documentation +(`a6e5bbf`), the ttaw 0.19.0 re-vendor with the `serve.env` pins of the numerics knobs, the quickstart picture and two +test changes (`d521981`), the ttaw 0.20.0 re-vendor (`c0d84f9`), and the release docs and demo media. After the re-vendor the device suite was re-run with +identical gate values and per-scene numbers (`VERIFICATION_2026-10-08.md`). + +## Round 1 + +Not started. The optimization phase follows the publication of every model of the collection (`research/PLAN.md` +§4.1, §5): each step is one commit with one `DIFFUSION_PLANNER_*` A/B knob, keeps the frozen gates and is re-checked on +all 99 end-to-end scenes (the plan's sensitivity to device numerics is chaotic per scene: PORT_LOG known issues). The +precision policy (PLAN.md §9.3) may be relaxed per module only with gate evidence; the precision cost is to be +recovered with fused kernels that keep fp32-level accuracy, not by dropping precision. + +### Findings (measured) + +The baseline profile (`OPT_BASELINE.md` "Device profile") is the starting point: + +- The trace is **kernel-bound, not dispatch-bound**: op-to-op gaps total 4.39 ms of a 103.66 ms span (median gap + 0.56 µs). Removing a short program saves its ~5-8 µs of kernel time plus ~0.6 µs. +- The plan does 165 GFLOP of matmul (split passes counted), ~1 ms at the 120-core HiFi4 peak: the time goes to badly + parallelised matmuls and thousands of small element-wise programs. Matmul efficiency is 7.2 % of the FLOP roofline + on average. +- By layer kind: split matmuls 34.8 ms, plain linears 24.0 ms (23.3 of them in the encoder's mixer trunks), fp32 + LayerNorm 21.6 ms, fp32 matmul attention 16.4 ms, glue 5.6 ms. +- What each numerics default costs (device time per plan, timing only; only the shipped configuration is gate-clean): + decoder split matmuls 20.2 ms, fp32 LayerNorm 18.1 ms, fp32 matmul attention 15.5 ms, encoder split matmuls 6.0 ms; + all four together 58.3 ms (the fastest graph, 43.74 ms, fails `enc.ego`), the decision-11 additions over the first + device round 33.4 ms (round-1 defaults: 68.64 ms, which fail `ego.mean_err_m` on nuScenes scene-0103_kf14). +- Dispatch / CQ matrix (back-to-back): ETH-1CQ 102.04, ETH-2CQ 111.87, WORKER-2CQ 101.88, WORKER-1CQ 101.87 ms. WORKER + is equal (the 12th column hardly matters to this graph yet), so ETH stays the default (D14); 2 CQs cost +9.8 ms on ETH, + so 1 CQ stays pinned. + +### Megakernel / fusion work + +None yet. The candidates (D19) are in the backlog (item 6): the persistent DiT-evaluation kernel first. + +## Known hangs (all rounds) + +| when (UTC) | command | cause | status | +|---|---|---|---| +| – | – | none: no hang, timeout, reset or FAULT marker in any device job of the port, the baseline or the release docs | – | + +## Rejected / not kept + +None yet. Configurations that were measured and are **not** the default, for accuracy (PORT_LOG decision 11, +`OPT_BASELINE.md` "What the numerics defaults cost"): + +- the first device round's defaults (split only for the mixer inputs and the decoder pre-projection, bf16 SDPA): + 68.7 ms, but ego mean 0.347 m > 0.3 m on nuScenes scene-0103_kf14 (job 11); +- configuration B (A + split and fp32 LayerNorm in the whole encoder): 109 ms, worst ego 0.35 / 0.15 m; +- configuration C (split and fp32 LayerNorm everywhere): 169 ms, worst ego 0.22 / 0.10 m: about a third less error + than the shipped configuration A (0.31 / 0.14 m, 105 ms in the same job 15) at +61 % device time (PORT_LOG open + question 8); +- the fastest graph (fused LayerNorm, no split, bf16 SDPA): 43.7 ms, fails the encoder PCC gate (`enc.ego` 0.99856 / + 0.99261). + +## Profile at the end (trace replay) + +The baseline profile (`OPT_BASELINE.md`): 6,282 programs, kernel sum 99.27 ms, span 103.66 ms; by op code Matmul +52.3 ms (1,365 programs), BinaryNg 31.5 ms (2,948), Reduce 3.7 ms, Softmax 3.5 ms, Unary 2.8 ms, Typecast 2.4 ms. The +ten slowest programs are the neighbour mixer's channel matmuls `[1, 320, 64, 128] @ [128, 128]`, ~970 µs each on 8 +cores. + +## Remaining backlog (gains on the 102.0 ms replay unless marked e2e; estimates, not measurements) + +From `OPT_BASELINE.md` "Ranked optimization opportunities". The gains are not additive: each item shrinks the base of +the next. + +1. **Encoder channel MLPs on 2-D activations (≈ −20 to −22 ms; exact rewrite, low effort).** The mixer `channels_mlp` + and the channel pre-projections run `ttnn.linear` on 4-D `[1, E, T, C]` activations, a batched matmul on 4-8 cores + (90 programs, 25.8 ms). A free `[1, 1, E·64, 128]` view for the MixerBlocks (as the token-mixing matmuls already do: + 69 µs on 117 cores) and a re-tiled 2-D layout or an explicit multi-core program config (with the HiFi4 compute + config: a `program_config` alone falls back to LoFi) for the pre-projections. +2. **A fused split (bf16x3) matmul, one program per linear (≈ −17 to −19 ms).** 338 split linears are 2,802 programs + and 34.8 ms today; a `generic_op` matmul that splits hi / lo on the fly and accumulates the three products in fp32 + DEST removes the typecast / subtract / add programs (9.3 ms in the decoder) and shares the operand reads. +3. **A fused fp32 LayerNorm, one program per LN (≈ −15 to −16 ms).** 237 LNs are 2,133 programs and 21.6 ms; the fused + `ttnn.layer_norm` is ~2.5e-3 relative and fails the mixers (PORT_LOG job 5), so the kernel must keep fp32 + statistics. +4. **Attention (≈ −5.5 ms with stock ops, ≈ −12 to −13 ms with a fused kernel).** 72 fp32 matmul attentions, 16.4 ms: + fold the 1/√32 scale into the Q weights and fuse the mask into the softmax (stock ops), spread the P·V matmuls over + heads, or one fp32-accurate flash-attention `generic_op` (scores in L1, hi / lo operands). +5. **Glue (≈ −3 to −4 ms).** Mixer transposes and residual adds, the adaLN gate multiplies (exact fold into per-step + weight copies), the solver update as one program instead of 3-5. +6. **Megakernels (D19; after 1-4).** MK-D, the persistent DiT-evaluation kernel (today 4.67 ms in ~435 programs for + 12.0 GFLOP; its weights fit in L1; target 0.5-1 ms per evaluation, the decoder from 51 to ~6-11 ms); MK-E / MK-F for + a mixer trunk / a fusion block; a single persistent megakernel for the whole plan is plausible (48 MB of weights and + constants, the largest activation 10.5 MB, ~175 MB of aggregate L1). It ships only if faster with the gates intact; + the attempt is recorded either way. +7. **Compaction (LOAD-time buckets; 20-40 % of what remains after 1-4, to be measured).** The device time is + scene-independent: kashiwanoha uses 88 / 320 neighbours and 89 / 352 decoder rows. Valid-entity buckets (one trace + per bucket, exact compaction rewrites) shrink the encoder in proportion and the attention scores quadratically. +8. **Host side (e2e ≈ −4 to −6 ms per request).** 15.8 ms of host work and transfers surround the replay: a 2-D + `neighbor_x` upload (1.31 MB of TILE padding -> 245 KB), skipping the `y0` upload while `x_T` is zero and sending + `cs` as its 4 columns (H2D 3.6 -> ~1.6 MB), vectorised per-entity loops in the pre- and post-processing. +9. **Dispatch configuration: no gain today; re-measure after 1-5.** Re-run the ETH / WORKER and 1 / 2 CQ matrix once + the program count drops below ~2,000; the ETH-2CQ penalty (+1.6 µs per program) needs a dispatch-core profile + first. diff --git a/PYTHON.md b/PYTHON.md new file mode 100644 index 0000000000000000000000000000000000000000..0d29e7c92256b84088f1eb229c061c46a9e10d0c --- /dev/null +++ b/PYTHON.md @@ -0,0 +1,179 @@ +# Python API: Diffusion Planner v5.0 (Autoware diffusion_planner) on Blackhole + +Use this API from Python code (a pipeline, a notebook, a ROS 2 node wrapper). You do not need the HTTP server: the +API and the server share the decoders, the device trace and the post-processing, so the outputs and the speed are +the same. + +## Install + +Install the package on top of an environment that already has `ttnn` (a tt-metal `python_env` at `44d66500520` +with `patches/tt-metal-eth-dispatch.patch`, or the tt-model container). From the root of the model repository (the +directory that holds `pyproject.toml`, `README.md` and `code/`): + +```bash +pip install -e . # the Python API (numpy<2, pillow, pyyaml, onnx, huggingface_hub) +pip install -e ".[server,test]" # + the HTTP server and the tests +``` + +The pip project is the repository's top-level `pyproject.toml`; it installs the package from +`code/tt_diffusion_planner` (there is no `pyproject.toml` inside `code/`, because the container build copies `code/` +over the tt-metal tree). ttnn and torch come from tt-metal and are not declared. + +The package carries `tt_diffusion_planner.ttaw`, the shared code of the Autoware ports to Blackhole (device open, trace +runner, decoders, model base class, HTTP app), vendored at the version recorded in +`code/tt_diffusion_planner/ttaw/VENDORED.json`. + +| You want to run | Extras | +|---|---| +| the Python API | none | +| the HTTP server (`tt_diffusion_planner.server.app`, see `SERVING.md`) | `server` | +| host tests (no device; device tests are skipped): `TT_VISIBLE_DEVICES=none python -m pytest -q code/tt_diffusion_planner/tests` | `server,test` | +| device tests: `python -m pytest -q -s code/tt_diffusion_planner/tests/test_pcc_device.py code/tt_diffusion_planner/tests/test_e2e_device.py` | `test` | + +## Quickstart + +```python +from tt_diffusion_planner import DiffusionPlanner + +with DiffusionPlanner.from_pretrained(device_id=0) as model: + out = model(inputs="code/tt_diffusion_planner/samples/kashiwanoha_dense.npz") +print(out.to_dict()) # the POST /predict body +``` + +`examples/quickstart.py` runs the same snippet, writes `quickstart.json` and a bird's-eye view of the input and the +plan (`quickstart_bev.png`). + +## `DiffusionPlanner.from_pretrained(...)` + +```python +DiffusionPlanner.from_pretrained( + model_id=None, # HF repo or a local directory with the weights files; default AutowareFoundation/diffusion_planner + *, + revision=None, # default for the default repo: the validated commit 423efde67f5 (tag v5.0) + variant=None, # "default" (the only v5.0 graph); default $DIFFUSION_PLANNER_VARIANT or "default" + device_id=None, # chip to open; default $TT_DEVICE_ID or 0 + device=None, # an already-opened ttnn device (tt_diffusion_planner.device.open_device); close() does not close it + dispatch=None, # "eth" (p150 target, 12x10 grid) | "worker" (A/B only, 11x10) | "auto"; default $DIFFUSION_PLANNER_DISPATCH or "eth" + num_command_queues=None, # default $DIFFUSION_PLANNER_NUM_CQS or 1 + weights_dir=None, # explicit local weights directory; no Hub access + warmup_variants="default", # trace variants to capture now; see "Warm-up" + verbose=False, + precision=None, # the only compile parameter: extra precision-policy rules, e.g. "dec.*=HiFi2+fp32" (experiments only) +) -> DiffusionPlanner +``` + +What it does: resolves the weights first, so a Hub problem never claims the chip (`weights_dir` > +`$DIFFUSION_PLANNER_WEIGHTS_DIR` > a local `model_id` directory > the HF snapshot at the pinned revision, restricted to +the three v5.0 ONNX files and `diffusion_planner.param.json`, with an offline fallback to the cache; the sha256 of every +file and the weights' `major_version == 5` are checked), opens the chip (ETH dispatch, 12×10, 1 CQ; the other open +parameters are `DEVICE_DEFAULTS` in `tt_diffusion_planner/device.py`, overridable with `DIFFUSION_PLANNER_*`), reads the +ONNX initializers as data, uploads the weights and constants (48.1 MB), builds the graph, then compiles and captures +the metal trace. If ETH dispatch cannot open (tt-metal without the patch), it warns and falls back to WORKER dispatch +(`model.info["device"]["fallback"]` names it). Any other keyword argument is a `TypeError`. + +The numerics are not arguments: the published configuration is the default of the `DIFFUSION_PLANNER_LN_FP32`, +`_SPLIT_MATMUL`, `_ATTN_MATMUL`, `_HIDDEN_FP32` and `_ATTN_FP32_ACC` knobs (`tt_diffusion_planner.tt.config.KNOBS`, +pinned in `tt-model.yaml` `serve.env`). Setting one of them in the environment changes the graph and invalidates the +accuracy figures of the card until the gates are re-run. + +## Warm-up + +`from_pretrained` returns a warm model: it builds the graph, runs the plan once eagerly (this first run compiles every kernel into the JIT cache), then captures the whole plan as one metal trace (`warmup_variants="default"`: the variant `plan`) with program-cache misses forbidden, so no later call compiles anything. `model.warmup()` is idempotent; `warmup_variants="none"` defers the capture to `model.warmup()`. + +Measured on the shipped sample (`model.info["warmup_ms"]`, 2026-10-09): the load takes 315 s with an empty JIT cache and 8.6 s with a warm one (build 0.53 s: the ONNX initializers read and 48.1 MB of weights and constants uploaded; warm-up and capture 3.7 s; the rest is the device open). The first call then takes 120 ms and the second 120 ms (the stage bench's steady state: 118 ms p50 for decoded arrays, 125 ms for an `.npz` path). The trace holds 74.6 MB of DRAM (`trace_region_size` 192 MiB). + +## Call: `model(...)` + +| Argument | Type | Description | +|---|---|---| +| `inputs` | mapping / `.npz` path / bytes / JSON envelope | the 15 raw planner tensors (see "Input types") | +| `velocity_smoothing_window` | int, 1..79, default 8 | forward moving average of the trajectory velocity, in points | +| `stopping_threshold` | float >= 0, default 0.3 | force stop below this smoothed speed (m/s), when the ego moves | +| `turn_indicator_keep_offset` | float, default -1.25 | added to the KEEP logit before the turn-indicator decision | +| `return_denoising_steps` | bool, default False | add the ego row of the 11 solver iterates (`out.meta["denoising_steps"]`, `[11, 81, 4]`, the node's `~/debug/denoising_steps`) | + +### Input types + +- `inputs=`: the 15 raw tensors of the Autoware node's `DiffusionPlannerCore::create_input_data()` (batch 1, + float32, ego `base_link` frame, BEFORE normalization; names and shapes in `tt_diffusion_planner.INPUT_SCHEMA`): + a `{name: array}` mapping (numpy or torch), an `.npz` path or its bytes, or the `/predict` envelope + `{"format": "npz", "data": }` / `{"format": "json", "arrays": {...}}`. Names, shapes and finite values are + checked (`InputError`). `tt_diffusion_planner.load_inputs(source)` is the same decoder. +- Any other input (`points`, `images`, `calibration`, ...) is refused (`InputError`). + +### What the caller keeps (the API is stateless) + +One call is one independent plan. The Autoware node keeps state between plans; to reproduce it over a sequence of +plans, the caller keeps the same state (SERVING.md 3.5 has the details): + +- **The tensors.** The node's pre-processing from ROS messages and the Lanelet2 map (per-UUID agent buffers and their + 0.1 s resampling, the ego history, lane / route / polygon / line-string selection and encoding, traffic lights, + speed limits, goal, turn-indicator report history) is not part of the bundle. +- **The turn-indicator hold window** (`turn_indicator_hold_duration`, 1.0 s in the node's YAML). Each call decides + with a fresh manager; apply the node's hold across calls with the node's own manager: + + ```python + from tt_diffusion_planner.host.postprocess import TurnIndicatorManager + + manager = TurnIndicatorManager() # hold 1.0 s, KEEP offset -1.25 (the node's YAML) + out = model(inputs=tensors) + decision = manager.evaluate(out.turn_indicator["logits"], stamp_s=now_s, prev_report=int(tensors["turn_indicators"][0, 30])) + command = decision.command # the held command while less than 1.0 s has passed + ``` + +- **The initial solver state** `sampled_trajectories` (`x_T`, normalised space): zeros is the node's default + (`temperature: [0.0]`); for a temperature > 0 send N(0, 1) x temperature; for the RTC prefix (`delay_step` > 0) put + the previous plan into the ego row, slots t = 0 .. delay_step (x as (x - 10) / 20, y as y / 20, cos / sin as they are, + in the current ego frame). `delay` is accepted and ignored (the node's multi-step mode never reads it). +- **The map frame.** Outputs are in `base_link`; the node transforms them with the current ego pose to `map`. + +## Output + +`model(...)` returns a `tt_diffusion_planner.Output` (= `ttaw.outputs.Trajectory`); `out.to_dict()` is exactly the `POST /predict` body (SERVING.md section 3.2). + +| field | type | meaning | +|---|---|---| +| `poses` | float32 `[80, 7]` | the ego trajectory at 0.1-8.0 s in `base_link`: x, y, yaw, cos, sin, velocity, acceleration (`out.columns`), post-processed like the node's `~/output/trajectory` | +| `turn_indicator` | dict | `command` (0 NO_COMMAND, 1 DISABLE, 2 ENABLE_LEFT, 3 ENABLE_RIGHT), `command_name`, `keep_selected`, `held` (always false: no hold window), the 5 raw `logits` (NONE, DISABLE, LEFT, RIGHT, KEEP), the decision's `probabilities` | +| `predicted_agents` | float32 `[N, 80, 5]` | x, y, yaw, cos, sin of each non-empty neighbour row, in input order | +| `meta` | dict | `predicted_agent_rows` (the rows of `predicted_agents`), `predicted_agent_columns`, `force_stop`, `time_from_start_s`, `valid_counts` (the entities the encoder saw), and with `return_denoising_steps` the encoded `denoising_steps` | +| `timing_ms` | dict | `preprocess`, `device` (host tensors + H2D + replay + D2H), `postprocess`, `total` | + +`out.to_dicts()` gives one `{x, y, yaw, cos, sin, velocity, acceleration}` dict per trajectory point; `out.to_dict("npz")` adds the poses as a lossless base64 NPZ array. + +## Lifetime and information + +- `model.close()` releases the trace and the persistent device tensors and closes the chip if the model opened it; + idempotent. `with` calls it for you; an unclosed model is closed when Python exits. +- `model.info`: weights (repo, tag, revision, path), device (dispatch, grid, CQs, fallback), variant, warm variants, + warm-up times, runtime parameter defaults, the input schema, the numerics options and precision policy in effect, the + trace (variants, persistent inputs, trace buffers in MB). +- Calls from several threads are safe: the device calls are serialised. One model per process per chip. + +## Speed + +Warm calls, batch 1, ETH dispatch, 1 CQ, 12×10, the pinned numerics (`code/scripts/bench.py`, 100 iterations; the numbers of `OPT_BASELINE.md`, 2026-10-08, on a shared host; p50, with p99 in brackets): + +| stage | shipped sample `kashiwanoha_dense` | +|---|---:| +| `.npz` decode + schema check (path inputs only) | 7.02 (18.20) ms | +| host pre-processing (the node's normalization, the encoder's host features, masks, `x_T`) | 4.75 (12.55) ms | +| packing the 19 persistent trace inputs · ttnn host tensors · H2D | 0.45 · 2.41 · 1.14 ms | +| **device trace, one blocking plan** | **102.13** (104.81) ms | +| D2H (one packed read, 460 KB) | 0.58 ms | +| host post-processing (trajectory, predicted paths, turn decision) | 4.32 (11.54) ms | +| **`model(inputs=arrays)` end to end** | **117.90** (134.56) ms | +| `model(inputs=<.npz path>)` | 124.76 (147.74) ms | +| back-to-back replays (device time per plan) | 102.04 ms = 9.80 plans/s | + +The device time does not depend on the scene: every plan computes the full capacities (re-checked on c0d84f9: kashiwanoha_dense 102.10, straight_road 102.11, a nuScenes instant 102.09 ms). Throughput above one plan per ~118 ms needs pipelining of the host work of neighbouring requests (not implemented); 2 CQs do not help a synchronous request (`OPT_BASELINE.md`). Where the time goes and what comes next: `OPT_REPORT.md`. + +## Limits + +- Batch 1 on the chip; one model per process. +- Fixed shapes of the v5.0 export (320 neighbours, 140 lanes, 25 route lanes, 10 polygons, 60 line strings, 31 history + and 80 future steps) and 10 DPM-Solver steps (11 decoder evaluations) are compiled into the trace; every plan computes + the full capacities, so the device time does not depend on the scene. +- The node's guidance services (start / stop / centerline guidance) are not available (the node's default is off). +- Accuracy is agreement with the fp32 CPU reference of the same network (README "Demo & Performances"); the planner's + driving quality is the weights' (trained by TIER IV on data that is not public). diff --git a/README.md b/README.md new file mode 100644 index 0000000000000000000000000000000000000000..b8b8cf16b3720c8fb86d0f9be31cc15c601abdc3 --- /dev/null +++ b/README.md @@ -0,0 +1,195 @@ +--- +tags: +- blackhole +- p150 +- tt-dit-server +- tt-model-cache +- tt-model-container +- tenstorrent +- ttnn +- tt-metal +- tt-nn +- autoware +- autonomous-driving +- planning +- trajectory-generation +- motion-prediction +- diffusion +- diffusion-planner +- tt-model-catalog +pipeline_tag: robotics +license: apache-2.0 +license_link: https://huggingface.co/AutowareFoundation/diffusion_planner +base_model: +- AutowareFoundation/diffusion_planner +--- + +# diffusion-planner-p150 + +Diffusion Planner v5.0 (Autoware diffusion_planner): the network Autoware deploys in `autoware_diffusion_planner`, ported to one Tenstorrent Blackhole p150 with tt-nn. The whole plan runs on the chip as one metal trace: the scene encoder, the 11 DiT decoder evaluations of the DPM-Solver++(2M) loop with their solver updates, and the turn-indicator head. The Autoware planner tensors in (ego and neighbour histories, lanes, route, polygons, line strings, goal, ego shape, turn-indicator history); an 8 s ego trajectory, the predicted 8 s paths of the neighbours and a turn-indicator command out, with the node's exact pre- and post-processing. +Weights: [AutowareFoundation/diffusion_planner `v5.0`](https://huggingface.co/AutowareFoundation/diffusion_planner/tree/423efde67f5414734da43a7ad856c17ceb8b51aa) · Paper: [arXiv:2501.15564](https://arxiv.org/abs/2501.15564) · Autoware package: [autoware_diffusion_planner](https://github.com/autowarefoundation/autoware_universe/tree/9ceaccf026c31ffc5319bc9eeb4bd7bede0af3fd/planning/autoware_diffusion_planner) · Training code: [tier4/Diffusion-Planner (TIER IV training fork)](https://github.com/tier4/Diffusion-Planner) · Port: [`code/`](https://huggingface.co/changh95/diffusion-planner-p150/tree/main/code) + +Runs on **p150** (mesh `P150`). Configuration: dispatch on the ETH cores, 1 command queue, 12×10 compute grid. Numerics (the default and the `serve.env` pins): fp32 residual streams and solver state, HiFi4 with fp32 accumulation, split hi / lo matmuls for the mixer inputs and every decoder linear, fp32 LayerNorm in the mixers and the decoder, fp32 matmul attention in the fusion encoder and the decoder: the configuration the end-to-end accuracy gates need (Caveats). All numbers on this card were measured in this configuration. + +Packaged and published with [tt-model-manager](https://github.com/tenstorrent/tt-model-manager) 0.1.0 (manifest schema 5.1). + +## Quickstart (Python) + +Prerequisite: a tt-metal / ttnn environment at tt-metal [`44d66500520`](https://github.com/tenstorrent/tt-metal/commit/44d66500520fda9f2c7060c0f6b41ec48f7ab37e) with [`patches/tt-metal-eth-dispatch.patch`](patches/tt-metal-eth-dispatch.patch) applied. ttnn is not on PyPI. + +```bash +hf download changh95/diffusion-planner-p150 --exclude "image/*" --local-dir diffusion-planner-p150 && cd diffusion-planner-p150 +pip install -e . # adds numpy<2, pillow, pyyaml, onnx, huggingface_hub; ttnn and torch come from tt-metal +pip install -e ".[server,test]" # optional: the HTTP server and the tests +``` + +Run the snippet from the model repo root: `code/tt_diffusion_planner/samples/kashiwanoha_dense.npz` is a path relative to it. + +```python +from tt_diffusion_planner import DiffusionPlanner + +with DiffusionPlanner.from_pretrained(device_id=0) as model: # weights -> your HF cache, trace captured + out = model(inputs="code/tt_diffusion_planner/samples/kashiwanoha_dense.npz") # the 15 raw planner tensors: .npz path, its bytes, or {name: array} + +print(out.columns) # x, y, yaw, cos, sin, velocity, acceleration (base_link, 0.1-8.0 s) +print(out.poses[:5]) +print(out.turn_indicator["command_name"], out.predicted_agents.shape) +``` + +- `from_pretrained` downloads the three v5.0 ONNX files and `diffusion_planner.param.json` (58.9 MB) of [`AutowareFoundation/diffusion_planner`](https://huggingface.co/AutowareFoundation/diffusion_planner) at the pinned commit `423efde67f5` (tag `v5.0`) to your HF cache (no token needed), checks their sha256, opens the chip, builds the graph and captures the metal trace. The first load compiles the kernels (315 s with an empty JIT cache, firmware and every kernel compiled); later loads take about 9 s (device open, reading and uploading the weights, warm-up and trace capture). +- The trace is captured during the load, so the first call is as fast as the later calls and no call compiles anything. +- The `with` block releases the trace and closes the chip. Without `with`, call `model.close()`. + +| | | +|---|---| +| **Input** | `inputs=`: the 15 raw tensors of the node's `create_input_data()` (ego frame, before normalization; batch 1): an `.npz` path or its bytes, a `{name: array}` mapping, or the `/predict` JSON envelope. Checked against `DiffusionPlanner.INPUT_SCHEMA` (names, shapes, finite values). Converting ROS messages and the Lanelet2 map into these tensors stays with the client. | +| **Options** | `velocity_smoothing_window=8`, `stopping_threshold=0.3`, `turn_indicator_keep_offset=-1.25`, `return_denoising_steps=False`. `from_pretrained(device_id=0, dispatch="eth", weights_dir=None, device=None)`. | +| **Output** | `Trajectory`: `poses` float32 [80, 7] (x, y, yaw, cos, sin, velocity, acceleration; base_link), `turn_indicator` (command, logits, probabilities), `predicted_agents` float32 [N, 80, 5], `timing_ms`, `meta` (the neighbour rows, force stop, the solver iterates on request). | +| **Methods** | `out.to_dict()` gives the `/predict` JSON. `out.to_dicts()` gives one dict per trajectory point. | + +- The API gives the same output as the HTTP server `/predict`: both share the decoders, the device trace and the host post-processing (checked on the device by `test_api_equals_server`). +- The API is stateless: one call is one plan. What the Autoware node keeps between plans stays with the caller: the turn-indicator hold window (1.0 s; the node's manager ships as `tt_diffusion_planner.host.postprocess.TurnIndicatorManager`), the initial solver state `sampled_trajectories` (zeros = the node's default temperature 0; noise for a temperature > 0; the previous plan for the RTC prefix) and the agent and ego histories. Details: [`code/PYTHON.md`](code/PYTHON.md) "What the caller keeps". +- One model uses one chip; calls from several threads are serialised. +- Full reference: [`code/PYTHON.md`](code/PYTHON.md). Runnable example: [`examples/quickstart.py`](examples/quickstart.py) (also writes `quickstart_bev.png`, the input tensors and the plan from above). + +## Serving (HTTP) + +```bash +tt-model pull changh95/diffusion-planner-p150 --with-weights +tt-model serve changh95/diffusion-planner-p150 # or with tt-cli: tt serve changh95/diffusion-planner-p150 +python3 code/tt_diffusion_planner/server/client.py --inputs code/tt_diffusion_planner/samples/kashiwanoha_dense.npz --out req.json +curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json +tt model stop changh95/diffusion-planner-p150 +``` + +- The image does not contain the weights. `--with-weights` puts them in your HF cache. +- The server uses port 20000 (or the next free port). It is ready when the log shows `Application startup complete`. +- One serve profile, the default (`tt-model profiles changh95/diffusion-planner-p150`). `serve.env` pins ETH dispatch, 1 CQ and the numerics knobs. +- Run the request lines from the model repo root (the `hf download` of the Quickstart): the client and the sample are files of this repo. `code/tt_diffusion_planner/server/client.py` needs only the Python standard library (no numpy); add `--url http://127.0.0.1:20000` to send the request. +- `POST /predict`: `inputs` (the 15 raw tensors of the Autoware node's `create_input_data()` in the ego frame, before normalization: a base64 `.npz` or `{"format": "json", "arrays": {...}}`); optional `params` (`velocity_smoothing_window` 8, `stopping_threshold` 0.3, `turn_indicator_keep_offset` -1.25, `return_denoising_steps` false), `output_format`. Also `GET /health`, `GET /info`, `GET /v1/models` (stub). Contract: [`SERVING.md`](SERVING.md) section 3. + +The response for the shipped sample (served on the p150; trajectory cut to 3 of 80 rows): + +```json +{ + "model": "diffusion-planner-p150", + "frame_id": "base_link", + "meta": {"predicted_agent_columns": ["x", "y", "yaw", "cos", "sin"], "predicted_agent_rows": [0, 1, 2, "... 85 more"], "force_stop": false, "time_from_start_s": [0.1, 0.2, 0.3, "... 77 more"], "valid_counts": {"ego": 1, "neighbor": 88, "static": 0, "lane": 123, "route": 17, "polygon": 0, "line_string": 60, "goal": 1, "ego_shape": 1, "turn": 1}}, + "timing_ms": {"preprocess": 3.8, "device": 105.3, "postprocess": 3.2, "total": 122.8, "decode": 8.5, "model_call": 112.5}, + "num_poses": 80, + "columns": ["x", "y", "yaw", "cos", "sin", "velocity", "acceleration"], + "trajectory": [[0.365, 0.01, -0.017, 0.9962, -0.0169, 3.8059, -0.0431], [0.7631, -0.0022, -0.0414, 0.9966, -0.0413, 3.8016, -0.5485], [1.1524, -0.0282, -0.0675, 0.9942, -0.0674, 3.7468, -0.4708], "... 77 more rows"], + "turn_indicator": {"command": 1, "command_name": "DISABLE", "keep_selected": true, "held": false, "logits": [-15.960084915161133, -5.125061511993408, -5.230083465576172, -0.7658059597015381, 5.504596710205078], "probabilities": [1.651601411190029e-09, 8.38487030705437e-05, 7.548934809165075e-05, 0.006556871347129345, 0.993184506893158]}, + "predicted_agents": {"format": "npz", "key": "predicted_agents", "dtype": "float32", "shape": [88, 80, 5], "data": ""} +} +``` + +- `trajectory`: 80 points at 0.1-8.0 s in `base_link` (the ego frame of the input tensors), post-processed like the node's `~/output/trajectory`: velocity from consecutive points, forward moving average over `velocity_smoothing_window` points, force stop (poses frozen once the smoothed speed falls below `stopping_threshold` while the ego moves), acceleration by finite difference. `yaw` is what `tf2::getYaw` reads from the node's (unnormalised) quaternion; `cos` / `sin` are the raw network outputs. +- `turn_indicator.command`: 0 NO_COMMAND, 1 DISABLE, 2 ENABLE_LEFT, 3 ENABLE_RIGHT (KEEP repeats the last input report); the node's 1 s hold window needs state across calls and is not applied (see above). +- `predicted_agents`: one 80-point path (x, y, yaw, cos, sin) per non-empty neighbour row, in input order (`meta.predicted_agent_rows`). + +## Demo + +| Shipped sample `kashiwanoha_dense.npz`: the p150 plan next to the fp32 CPU reference | The 11 solver iterates of the same plan (ego row), p150 and CPU | +|:---:|:---:| +| ![](media/dp_kashiwanoha_dense_tt_vs_cpu.png) | ![](media/dp_kashiwanoha_dense_denoising_steps_tt.png) | +| **Shipped sample `straight_road.npz`** | **Agreement with the fp32 CPU reference on all 99 gated scenes** | +| ![](media/dp_straight_road_tt_vs_cpu.png) | ![](media/dp_tt_vs_cpu_agreement_99_scenes.png) | + +On nuScenes v1.0-mini planning instants (p150 outputs; the scenes are converted from the dataset into the planner tensors and are not in this repository). **Non-commercial, CC BY-NC-SA 4.0.** The grey path with hollow dots is the logged drive; the model never saw nuScenes (see Caveats). + +| scene-0061, 2 Hz sequence (33 plans): following a van, then a left turn | scene-0061 key-frame 18, in the turn: CAM_FRONT with the plan as a vehicle-width ribbon | +|:---:|:---:| +| ![](media/dp_nuscenes_scene-0061_bev_tt_NC.gif) | ![](media/dp_nuscenes_scene-0061_kf18_cam_front_tt_NC.jpg) | +| **scene-0757 key-frame 11: the turn head commands RIGHT before the driver's blinker** | **scene-0757 key-frame 11, CAM_FRONT** | +| ![](media/dp_nuscenes_scene-0757_kf11_bev_tt_NC.png) | ![](media/dp_nuscenes_scene-0757_kf11_cam_front_tt_NC.jpg) | +| **scene-0103 key-frame 12 (mini_val, Boston)** | **scene-0103 key-frame 12, CAM_FRONT** | +| ![](media/dp_nuscenes_scene-0103_kf12_bev_tt_NC.png) | ![](media/dp_nuscenes_scene-0103_kf12_cam_front_tt_NC.jpg) | +| **scene-0061 key-frame 6: approach behind the van** | **scene-0916 key-frame 11 (mini_val): parking-lot right turn** | +| ![](media/dp_nuscenes_scene-0061_kf06_bev_tt_NC.png) | ![](media/dp_nuscenes_scene-0916_kf11_bev_tt_NC.png) | + +nuScenes renders: rendered from the nuScenes dataset (v1.0-mini, CAN bus expansion and map expansion v1.3), © Motional AD Inc., CC BY-NC-SA 4.0 and the nuScenes Terms of Use; non-commercial use only; Motional does not endorse this work. The kashiwanoha sample is derived from the Apache-2.0 map AutowareFoundation/map-carla-kashiwanoha. Sources, changes and the full attributions: [`media/ATTRIBUTION.md`](media/ATTRIBUTION.md). + +## Demo & Performances + +Warm, batch 1; the stage bench 2026-10-08, the served rows, load times and the re-check 2026-10-09. Latency: the stage bench of [`OPT_BASELINE.md`](OPT_BASELINE.md) (`code/scripts/bench.py`, 100 iterations per stage) on the shipped sample `kashiwanoha_dense.npz` (88 neighbours, 123 lanes, 17 route lanes, 60 line strings; the device time is the same for every scene: every plan computes the full capacities); the served rows from uvicorn on the host (the app the container runs, with its serve pins) and a loopback client, 50 requests of the shipped sample. The host is shared with other jobs, so the host stages move with its load by several ms; the device rows repeat to ±0.05 ms. Accuracy: the p150 output against the fp32 CPU reference of the same network on 99 scenes: the 2 shipped samples, 5 research scenes and 92 nuScenes v1.0-mini planning instants (the frozen end-to-end gates), and against the research pipeline's ONNX Runtime outputs as an independent oracle. + +| Metric | Performance | +|---|---:| +| Agreement with the fp32 CPU reference, shipped sample `kashiwanoha_dense` (8 s ego plan, 88 neighbours) | ego max **2.3 cm** / mean 1.1 cm; turn command identical; neighbours: median per-agent max 8.6 cm | +| Agreement, shipped sample `straight_road` | ego max 4.9 cm / mean 1.8 cm; turn command identical; neighbours 6.3 cm | +| Agreement, all 99 gated scenes (2 samples, 5 research scenes, 92 nuScenes-mini instants) | worst ego max **0.313 m** / mean **0.143 m** (gates 1.0 / 0.3 m; nuscenes/scene-0103_kf14); ego mean: median 1.3 cm, 95th percentile 6.2 cm; turn command identical **99 / 99**; neighbours: median per-agent max ≤ 0.086 m (gate 1.5 m) | +| Agreement with an independent oracle: the research pipeline's ONNX Runtime outputs (raw x0), 92 nuScenes instants + the 33-plan scene-0061 sequence | worst ego max 0.313 m / mean 0.143 m, turn command identical 92 / 92; sequence: worst ego max 0.118 m / mean 0.049 m, turn 33 / 33; every instant within the gates | +| Module PCC vs the fp32 reference (encoder categories, encoding, teacher-forced decoder evaluation; gate 0.999) | ≥ 0.999952 / 0.999978 / ≥ 0.9999995 | +| Open-loop vs the nuScenes log, 92 instants (a sanity check against one human driver, not a planning metric) | p150 ADE / FDE at 8 s 4.75 / 11.87 m (CPU reference 4.74 / 11.87; constant velocity 4.99 / 13.06); turn command = logged blinker 88 / 92 (CPU 88 / 92) | +| Python `model()` call, shipped sample (host pre-processing, H2D, trace, D2H, host post-processing) | **117.9 ms p50** (p99 134.6) · 8.5 plans/s | +| Served `/predict` `timing_ms.total` (uvicorn on the host, the shipped sample) | **123.1 ms median** (min 122.5; of which decode 8.6) | +| Served client round trip, loopback (base64 `.npz` request, 0.15 MB) | 127.5 ms median | +| Device trace, one blocking plan (encoder + 11 DiT evaluations + 10 solver updates + turn head) | **102.13 ms** | +| Back-to-back trace replays | **102.04 ms per plan** · 9.80 plans/s | +| Host pre-processing · pack · host tensors · H2D · D2H · host post-processing | 4.75 · 0.45 · 2.41 · 1.14 · 0.58 · 4.32 ms | +| `from_pretrained` load: empty JIT cache / warm cache | 315 s / 8.6 s | + +All numbers in this table were measured with dispatch on the ETH cores, 1 command queue and a 12×10 compute grid on one p150, with the pinned numerics of `serve.env`. Accuracy is agreement with the fp32 CPU reference of the same Autoware network (same weights, same pre- and post-processing); no dataset-level accuracy is claimed: the paper's benchmark is nuPlan closed loop (an account-gated dataset and simulator, not run here), and the deployed v5.0 weights were trained by TIER IV on data that is not public, so no public benchmark is in-domain. Details: [`VERIFICATION_2026-10-08.md`](VERIFICATION_2026-10-08.md), [`OPT_BASELINE.md`](OPT_BASELINE.md), [`OPT_REPORT.md`](OPT_REPORT.md). + +No GPU comparison: no GPU was available on the host where this port was built and measured, so this card makes no GPU speed claim. The reference rows are the port's own fp32 CPU reference on the same host (a correctness baseline, not a speed target). Autoware's CHANGELOG quotes 5.13 ms mean (300 runs) for an older single-step engine of this planner on an RTX PRO 6000 Blackwell with TensorRT (precision not stated); it is not like-for-like with this v5.0 multi-step port. p150 power was not measured, so no efficiency comparison is made. + +## Caveats + +- First release: **baseline port, optimization pending.** The plan is one metal trace and kernel-bound (6,282 programs; op-to-op gaps 4.4 ms of a 103.7 ms span). The **first optimization target is the precision cost**: the numerics defaults that the end-to-end gates need (split hi / lo matmuls, fp32 LayerNorm, fp32 matmul attention) cost 33.4 ms per plan (102.0 ms vs 68.6 ms with the first device round's defaults, which fail the gates); fused kernels are to recover it without dropping precision. A second sink of similar size is unrelated to precision: the encoder's channel-MLP and pre-projection matmuls run on 4-8 cores (25.8 ms). [`OPT_REPORT.md`](OPT_REPORT.md) ranks what comes next. +- Deployment status in Autoware: `autoware_diffusion_planner` is an alternative to the default rule-based planning stack, selected with `planning_setting:=diffusion_planner` (package README); it is aimed at Autoware's proposed new planning framework. This bundle is not a ROS 2 node (Python API and HTTP) and not a certified Autoware component; do not use it for safety-critical driving decisions or closed-loop vehicle control. +- Stateless API: the node's state between plans (the turn-indicator hold window, the RTC prefix and temperature of the initial solver state, the agent buffers and the ego history) is the client's ("What the caller keeps" in [`code/PYTHON.md`](code/PYTHON.md)). The node's guidance services (start / stop / centerline guidance) are off, as in the node's default. +- Precision policy of this release: fp32 residual streams and solver state; HiFi4 with fp32 accumulation for every matmul; the ego / neighbour pre-projection as a pad-relative fp32 island; split hi / lo matmuls (bf16 hi + fp32 lo parts, ~1e-5 relative, because a device fp32 matmul rounds its operands like TF32) for the mixer inputs and every decoder linear; an fp32 LayerNorm decomposition in the mixers and the decoder (the fused `ttnn.layer_norm` loses the per-entity signal on the mixers' offset-dominated rows); fp32 matmul attention in the fusion encoder and the decoder; the turn head in fp32; the other weights and the hidden MLP activations of the mixers and the fusion encoder in bf16. The plan's sensitivity to these choices is chaotic per scene: with the first device round's numerics, one nuScenes instant (scene-0103_kf14) moved 0.35 m on average while the module PCCs differed only in the 5th decimal. So every numerics change is re-checked on all 99 scenes. +- Validation scope: the p150 output agrees with the fp32 CPU reference on 99 scenes (above). The two shipped samples and the five research scenes are synthetic-but-faithful scenes built with a Python port of the node's tensor construction; the 92 nuScenes instants are converted from nuScenes v1.0-mini, a domain the model never saw (Singapore and Boston, right-hand traffic in Boston, oracle tracks from 2 Hz annotations, no traffic-light states, no speed limits, stop areas instead of stop lines). On them the plans are plausible but conservative (moving plans about 17 % shorter than the logged drive); the open-loop numbers above are a sanity check against one human driver, not a planning metric. +- Neighbour predictions: the gate is the median over agents of each agent's max displacement, so single agents can differ more. On the shipped sample `kashiwanoha_dense` (container smoke, served p150 output vs the stored CPU reference) the per-agent max displacement is 8.6 cm median and 0.26 m at the 90th percentile; the worst of the 88 agents differs by 4.31 m. The ego plan is gated on its max and mean; neighbour paths only on that median. +- Documented deviations from the node: the output stays in `base_link` (the node transforms it to `map` with the ego pose); `yaw` is what `tf2::getYaw` reads from the node's quaternion of the unnormalised cos / sin rotation (it differs from atan2(sin, cos) when |(cos, sin)| ≠ 1, as in the node); the speed masks follow the node's TensorRT path (`> FLT_EPSILON`); the turn indicator is decided without the hold window; `delay` is accepted and ignored (the node's multi-step mode never reads it). +- Does not scale to multiple p150 in a mesh configuration. The build uses a 12×10 compute grid of Tensix cores: the dispatch functions move from one Tensix column to the ETH cores (`patches/tt-metal-eth-dispatch.patch`), so this build assumes that you do not need chip-to-chip ethernet communication. +- `dispatch="worker"` (server: `DIFFUSION_PLANNER_DISPATCH=worker`) is an A/B opt-in. On a p150 it gives an 11×10 grid (101.87 ms per replay, the same as ETH within 0.2 %); if ETH dispatch is not available (tt-metal without the patch), the model falls back to it with a warning. The numbers on this card do not apply to that mode. +- Batch 1, one plan per request; requests are serialised on the chip. The shapes of the v5.0 export (320 neighbours, 140 lanes, 25 route lanes, 10 polygons, 60 line strings, 31 history and 80 future steps) and 10 DPM-Solver steps are compiled into the trace, and every plan computes them in full, so the device time does not depend on the scene. +- Not an OpenAI-compatible API; `GET /v1/models` is a stub so the tt-model ready card does not 404. +- p150 power was not measured, so no efficiency comparison is made. + +## Licensing + +- Weights: [AutowareFoundation/diffusion_planner](https://huggingface.co/AutowareFoundation/diffusion_planner) at tag `v5.0` (commit `423efde67f5414734da43a7ad856c17ceb8b51aa`), Apache-2.0 per its model card. Not redistributed here: the package only points to them. The upstream card states that TIER IV trained the models on TIER IV synthetic and real driving data; the dataset composition is not publicly documented. +- Pre- and post-processing ported from autoware_universe [`planning/autoware_diffusion_planner`](https://github.com/autowarefoundation/autoware_universe/tree/9ceaccf026c31ffc5319bc9eeb4bd7bede0af3fd/planning/autoware_diffusion_planner) (Apache-2.0). +- Port and serving code (`code/`): Apache-2.0. `patches/tt-metal-eth-dispatch.patch` modifies tt-metal (Apache-2.0). +- Sample data: `code/tt_diffusion_planner/samples/kashiwanoha_dense.npz` is derived from the Apache-2.0 Lanelet2 map [AutowareFoundation/map-carla-kashiwanoha](https://huggingface.co/datasets/AutowareFoundation/map-carla-kashiwanoha) (0.2.0) with a scripted ego, route and agents; `straight_road.npz` is a procedural scene generated by this repo. Both Apache-2.0, each with its stored CPU-reference output (`*.reference.json`). Only these redistributable samples ship; the nuScenes-derived planning instants of the accuracy tables are not in this repository. +- Demo media (`media/`, sources and changes in [`media/ATTRIBUTION.md`](media/ATTRIBUTION.md)): + - nuScenes renders (`media/*_NC.*`), **non-commercial, CC BY-NC-SA 4.0**: Rendered from the nuScenes dataset (v1.0-mini, CAN bus expansion and map expansion v1.3), © Motional AD Inc., CC BY-NC-SA 4.0 and the nuScenes Terms of Use (https://www.nuscenes.org/terms-of-use). Non-commercial use only; adaptations under the same license. Motional does not endorse this work. Cite: H. Caesar et al., *nuScenes: A Multimodal Dataset for Autonomous Driving*, CVPR 2020. + - Renders of the shipped samples: Apache-2.0 (kashiwanoha map: AutowareFoundation/map-carla-kashiwanoha@0.2.0, Apache-2.0). + +## Provenance + +These are the exact sources the container image was built from: + +| component | built from | +| --- | --- | +| tt-metal | [`44d66500520fda9f2c7060c0f6b41ec48f7ab37e`](https://github.com/tenstorrent/tt-metal/commit/44d66500520fda9f2c7060c0f6b41ec48f7ab37e) + [`patches/tt-metal-eth-dispatch.patch`](patches/tt-metal-eth-dispatch.patch) (sha256 `08d0ddf6…45cc`, 4 files; dirty tree: the image includes the patch) | +| weights | [`AutowareFoundation/diffusion_planner@423efde67f5414734da43a7ad856c17ceb8b51aa`](https://huggingface.co/AutowareFoundation/diffusion_planner/tree/423efde67f5414734da43a7ad856c17ceb8b51aa) (tag `v5.0`), files `diffusion_planner_encoder.onnx, diffusion_planner_decoder.onnx, diffusion_planner_turn_indicator.onnx, diffusion_planner.param.json` (sha256 `2856886a…ca49`, `eb30c0c0…57ca`, `07acfb58…a732`, `ee3145b6…a268`, checked at load) | +| Autoware reference | autoware_universe [`9ceaccf026c31ffc5319bc9eeb4bd7bede0af3fd`](https://github.com/autowarefoundation/autoware_universe/tree/9ceaccf026c31ffc5319bc9eeb4bd7bede0af3fd) (`planning/autoware_diffusion_planner`, package 0.53.0, `multi_step` mode) | +| shared package | `ttaw` 0.20.0, vendored as `code/tt_diffusion_planner/ttaw` from the Autoware ports' shared `common` repository at commit `89dec49` (`code/tt_diffusion_planner/ttaw/VENDORED.json`: version, commit and per-file sha256) | +| `code/` digest (image) | `c0e7abb7888098a9` (sha256, first 16 hex digits; `built.code_sha256` of `tt_kernel_manifest.json`) | +| image | `tt-model/diffusion-planner-p150:3b96d8ea7190` (`sha256:3b96d8ea71902fe6a00f1792dd41290839b7758070431464cd613cde3f6bf909`) | +| base images | build stage `ghcr.io/tenstorrent/tt-metal/tt-metalium/ubuntu-22.04-dev-amd64:latest` @ `sha256:df9d279c7f85c17c6fad982d196802682d669cca1b7ced9cbaad8181339cd5fc`; runtime stage `docker.io/library/ubuntu:22.04` @ `sha256:5ec03bb3441e8b0bf3b4f9cd4629a1ae763010dc3035bb8da3ae6cf026486401` (tt-model's `FROM` tags float; these are the digests this build resolved, see [`build_info.json`](build_info.json)) | +| built | 2026-10-09T04:37:19+00:00 by tt-model 0.1.0 | diff --git a/SERVING.md b/SERVING.md new file mode 100644 index 0000000000000000000000000000000000000000..bd8a02b674975e0345fa0d84a9348503b7ef24da --- /dev/null +++ b/SERVING.md @@ -0,0 +1,242 @@ +# Serving Diffusion Planner v5.0 (Autoware diffusion_planner) on Blackhole with tt-model-manager + +This repo is the authoring source of the **tt-model container package** `changh95/diffusion-planner-p150` +(`kind: tt-dit-server`, schema 5.1): `tt-model.yaml` is the manifest, `code/` the port, and +`code/tt_diffusion_planner/server/app.py` the ASGI app uvicorn runs. Weights are a pinned pointer, never baked into the +image. The device open, the decoders, the Python-API contract and the HTTP contract come from +`code/tt_diffusion_planner/ttaw/`, the shared package of the Autoware ports vendored into this repo +(`ttaw/VENDORED.json` records its version, source commit and file hashes; it is not edited here). + +| item | value | +|---|---| +| tt-metal tree | `/home/ubuntu/experiments/tt-models/tt-metal` (main `44d66500520`, `v0.80.0-dev20261006-78-g44d6650052`, + `patches/tt-metal-eth-dispatch.patch`; torch 2.11.0+cpu) | +| weights | `AutowareFoundation/diffusion_planner` @ `423efde67f5414734da43a7ad856c17ceb8b51aa` (tag `v5.0`; files `diffusion_planner_encoder.onnx`, `diffusion_planner_decoder.onnx`, `diffusion_planner_turn_indicator.onnx`, `diffusion_planner.param.json`; 58.9 MB; Apache-2.0; public, ungated; sha256 of every file checked at load) | +| Autoware reference | `planning/autoware_diffusion_planner` @ autoware_universe `9ceaccf026c31ffc5319bc9eeb4bd7bede0af3fd` (package 0.53.0), the node's `multi_step` mode | +| shared package | `ttaw` 0.20.0 (`common` @ `89dec49`) | +| app | `tt_diffusion_planner.server.app:app` (uvicorn; the ASGI lifespan does weights -> device -> graph -> trace capture) | +| device recipe | `ttnn.open_device(device_id, dispatch_core_config=DispatchCoreConfig(ETH), l1_small_size=32768, trace_region_size=192 MiB, num_command_queues=1)` (`DEVICE_DEFAULTS` in `code/tt_diffusion_planner/device.py`; the plan's trace holds 74.6 MB) | +| graph | ONE metal trace per plan: the encoder (6 MLP-Mixer trunks, small encoders, 6 fusion blocks), the cross-attention K / V hoisted once, 11 DiT evaluations with the 10 fp32 DPM-Solver++(2M) updates and the prefix constraint, the turn-indicator head, one packed readback; 6,282 device programs | +| hardware | one Blackhole p150 (`hardware: p150`, `mesh_device: P150`, `TT_MESH_SHAPE=1x1`), 12×10 compute grid | + +## 1. Run on the HOST (hardware validation, no Docker) + +The tt-metal `python_env` has ttnn and torch; `bin/tt-env.sh` puts the workspace's dev overlay (fastapi, uvicorn, onnx) +on `PYTHONPATH` instead of installing into it. On the shared workspace box every command that opens the chip goes +through the device lock (`bin/devrun`). + +```bash +ROOT=/home/ubuntu/experiments/tt-models +source $ROOT/bin/tt-env.sh # TT_METAL_HOME, PYTHONPATH (+ the dev overlay: fastapi, uvicorn, onnx), python_env +cd $ROOT/bundles/diffusion-planner-p150 +export PYTHONPATH=$PWD/code:$PYTHONPATH +export HF_MODEL=AutowareFoundation/diffusion_planner TT_MODEL_WEIGHTS_REVISION=423efde67f5414734da43a7ad856c17ceb8b51aa TT_MESH_SHAPE=1x1 + +# host tests (no device; the device tests are skipped; ttnn is the fake of common/tests/host: a real ttnn host +# tensor opens the chip, so only devrun jobs may build one) +PYTHONPATH=$ROOT/research/packaging/scripts:$PYTHONPATH TT_VISIBLE_DEVICES=none \ + python -m pytest -q -p no:cacheprovider -p fake_ttnn_plugin code/tt_diffusion_planner/tests +# device tests (per-module PCC + end-to-end agreement; 99 scenes with the workspace goldens, the 2 shipped samples elsewhere) +$ROOT/bin/devrun -t 1800 -- python -m pytest -q -s code/tt_diffusion_planner/tests/test_pcc_device.py code/tt_diffusion_planner/tests/test_e2e_device.py +# the server, with the serve pins of tt-model.yaml, and the container smoke test against it +$ROOT/bin/devrun -t 1800 -- bash -c 'env DIFFUSION_PLANNER_DISPATCH=eth DIFFUSION_PLANNER_NUM_CQS=1 \ + python -m uvicorn --host 127.0.0.1 --port 20000 --lifespan on tt_diffusion_planner.server.app:app & \ + S=$!; python3 code/tt_diffusion_planner/server/smoke_test.py --url http://127.0.0.1:20000 --wait 1200; R=$?; kill -TERM $S; wait $S; exit $R' +``` + +Boot log landmarks (they drive `tt-model serve`'s progress view, `boot_progress.py` TT_DIT_PHASES): +`Loading weights` -> `Opening device` -> `Warming up: capturing trace ...` -> `Warmup complete` -> uvicorn +`Application startup complete`. A cold JIT cache compiles every kernel first (315 s measured for +`from_pretrained`); later boots take about 10 s to READY (measured on the host, warm JIT cache). Startup failures raise and uvicorn exits non-zero (no CPU +fallback). SIGTERM: the lifespan releases the trace and closes the chip (within `tt-model stop`'s 120 s budget). + +Offline override (no Hub): `DIFFUSION_PLANNER_WEIGHTS_DIR=`. + +## 2. Package, serve, push (Docker) + +Docker 29 on this box is **rootful** (`/var/lib/docker`), and the user's session may not carry the `docker` group: +wrap every docker-using command in `sg docker -c "..."` (no `docker-env.sh` is needed here; the reference bundles' +`docker-env.sh` was the authors' rootless-Docker box). Run every `tt-model` command **from this directory** +(`source.tt_metal` and `extra_code[].root: code` are CWD-relative) and point `--out` outside it. + +```bash +ROOT=/home/ubuntu/experiments/tt-models; cd $ROOT/bundles/diffusion-planner-p150 +TTM=$HOME/.local/share/uv/tools/tt-model/bin/python + +# offline validation (seconds; no docker, no device) -- must print VALID +$TTM -c "from tt_kernel.container_manifest import load_container_manifest as L; m=L('tt-model.yaml', check_sources=True); p=m.resolve_profile(); print('VALID', m.name, m.kind, p.hardware, p.mesh_device, m.weights_ref)" +python3 $ROOT/research/packaging/scripts/check_bundle.py . # conventions (add --stage publish before a push) + +# build (one at a time on the box: the absolute lock file): ~35 min cold, ~12 min after a code/ change (the C++ +# build re-runs, ccache-warm, whenever code/ changes), ~1 min for a manifest- or lock-only change +flock $ROOT/.package.lock sg docker -c "tt-model package --container tt-model.yaml --out $ROOT/build" # log: ~/.cache/tt-model/build/diffusion-planner-p150.log +# provenance: the base-image digests this build resolved (tt-model's FROM tags float) -> build_info.json (published) +sg docker -c "python3 $ROOT/research/packaging/scripts/record_build_info.py --staged $ROOT/build/diffusion-planner-p150 --bundle ." + +# serve + smoke + stop, all inside ONE device-lock window (there is one serve profile, the default; serve returns +# once READY and leaves the container running; the script's trap saves the container log and always runs +# `tt-model stop`, also when devrun's timeout fires: -k 150 > the 120 s grace). No token reaches the public weights +# download. Evidence: logs/smoke/ (--log-dir). +$ROOT/bin/devrun -t 3600 -k 150 -- env -u HF_TOKEN -u HUGGING_FACE_HUB_TOKEN \ + sg docker -c "bash code/scripts/container_smoke.sh $ROOT/build/diffusion-planner-p150 20000" + +# publish: fill the provenance from the staged manifest, run the publish gate, overlay the hand-written docs, the pip +# project and the build record onto the staged dir, then push (HF_TOKEN only from the environment) +python3 $ROOT/research/packaging/scripts/instantiate_bundle.py --values $ROOT/research/diffusion-planner/bundle_values.json \ + --fill . --manifest $ROOT/build/diffusion-planner-p150/tt_kernel_manifest.json +python3 $ROOT/research/packaging/scripts/check_bundle.py . --stage publish +rsync -a README.md SERVING.md OPT_BASELINE.md OPT_REPORT.md VERIFICATION_*.md tt-model.yaml pyproject.toml build_info.json \ + examples media patches $ROOT/build/diffusion-planner-p150/ +cp code/PYTHON.md $ROOT/build/diffusion-planner-p150/PYTHON.md +tt-model push $ROOT/build/diffusion-planner-p150 --publish +``` + +The smoke test fails unless `/info` reports ETH dispatch, the 12x10 grid and what the staged package pins +(`DIFFUSION_PLANNER_DISPATCH`, `DIFFUSION_PLANNER_NUM_CQS`, `DIFFUSION_PLANNER_VARIANT`, the weights revision), and it +compares the served trajectory of the shipped sample with the stored CPU-reference output +(`code/tt_diffusion_planner/samples/kashiwanoha_dense.reference.json`: average displacement <= 0.3 m, final +displacement <= 1.0 m). On the host, the same smoke test against uvicorn with the serve pins passed: +`PASS diffusion-planner-p150: profile=- variant=default dispatch=eth grid=12x10 cqs=1 n=80 turn=DISABLE reference=kashiwanoha_dense.reference.json (trajectory ade=0.0109 fde=0.0173; predicted_agents max_abs=6.283) device_ms=109.402 total_ms=133.384 rtt_ms=139.0` + +(`predicted_agents max_abs` is reported, not gated. Here it is the yaw column at ±π: -3.1413 vs +3.1415 +rad, the same heading. The neighbour paths are gated in the end-to-end tests, on the per-agent maximum displacement.) + +`serve` publishes the first free port from 20000 and exports into the container: `HF_MODEL=AutowareFoundation/diffusion_planner`, +`MESH_DEVICE=P150`, `TT_MESH_SHAPE=1x1`, `TT_MODEL_WEIGHTS_REVISION=423efde67f5414734da43a7ad856c17ceb8b51aa`, then `serve.env`. The HF cache is +mounted at `/hf` (rw), the JIT cache at `/cache` (host `~/.cache/tt-model/diffusion-planner-p150/cache`), and the container +sees only its own chip (`--device /dev/tenstorrent/`). tt-cli users: `tt serve changh95/diffusion-planner-p150` / +`tt model stop changh95/diffusion-planner-p150` (point tt at this tt-model with +`tt config set tools.override.tt-model ~/.local/bin/tt-model` or `TT_TOOL_BIN_TT_MODEL`). + +What `push` does to the repo: `code/` and `image/` on the Hub become exactly the staged trees (`extra_code.paths` + +`models/common/lightweightmodule.py`); every top-level file of the staged dir is uploaded (the overlaid README +replaces the generated card; `pyproject.toml`, `PYTHON.md` and `build_info.json` arrive the same way, never through +`code/`); top-level files already on the Hub and not in the staged dir are kept. + +## 3. The request / response contract + +| route | purpose | +|---|---| +| `GET /health` | `{"status": "ok" \| "starting" \| "error", "model", "device": {"dispatch", "grid", "cores", "device_id"}, "error"}`, always 200; `ok` only after warm-up | +| `GET /v1/health` | same as `/health` (tt-model's hint for non-chat packages points here) | +| `GET /info` | model, task, io, Autoware package + commit, weights {repo, tag, revision, path, license}, device (dispatch, grid, CQs, fallback), the input schema (the 15 tensors with shapes and dtypes), labels (turn-indicator logit order), variant, warm variants, runtime params + defaults, the numerics options and precision policy in effect, the trace (variants, persistent inputs, trace buffers), warm-up and boot times | +| `GET /v1/models` | `{"object": "list", "data": [{"id": "AutowareFoundation/diffusion_planner", "object": "model", "owned_by": "changh95"}]}` -- a stub so OpenAI-shaped probes do not 404; NOT a chat API | +| `POST /predict` | one plan -> the model output (below) | + +### 3.1 Request (`application/json`; unknown fields are a 422) + +| field | type | meaning | +|---|---|---| +| `inputs` | object | the planner tensors: `{"format": "npz", "data": }` or `{"format": "json", "arrays": {name: nested list}}` with exactly the 15 raw tensors of the node's `create_input_data()`: `sampled_trajectories` [1,321,81,4], `ego_agent_past` [1,31,4], `ego_current_state` [1,10], `neighbor_agents_past` [1,320,31,11], `static_objects` [1,5,10], `lanes` [1,140,20,33], `lanes_speed_limit` [1,140,1], `route_lanes` [1,25,20,33], `route_lanes_speed_limit` [1,25,1], `polygons` [1,10,40,3], `line_strings` [1,60,20,4], `goal_pose` [1,4], `ego_shape` [1,3], `turn_indicators` [1,31], `delay` [1,1]; float32, ego (`base_link`) frame, BEFORE normalization; names, shapes and finite values are checked (400) | +| `params` | object | per-request knobs (host post-processing only): `velocity_smoothing_window` (8, 1..79), `stopping_threshold` (0.3 m/s), `turn_indicator_keep_offset` (-1.25), `return_denoising_steps` (false) | +| `output_format` | str | `json` (default) or `npz` (adds the poses as a lossless base64 NPZ array) | + +Base64 may be standard or URL-safe, wrapped, with or without a `data:` prefix. `python3 +code/tt_diffusion_planner/server/client.py --inputs [--param name=value] --out req.json` builds the request +with the standard library only (no numpy); `--url http://127.0.0.1:20000` sends it. + +### 3.2 Response (200) + +The served body of the shipped sample (uvicorn on the host with the serve pins; trajectory cut to 3 of 80 rows and +the predicted-agents array to its header): + +```json +{ + "model": "diffusion-planner-p150", + "frame_id": "base_link", + "meta": {"predicted_agent_columns": ["x", "y", "yaw", "cos", "sin"], "predicted_agent_rows": [0, 1, 2, "... 85 more"], "force_stop": false, "time_from_start_s": [0.1, 0.2, 0.3, "... 77 more"], "valid_counts": {"ego": 1, "neighbor": 88, "static": 0, "lane": 123, "route": 17, "polygon": 0, "line_string": 60, "goal": 1, "ego_shape": 1, "turn": 1}}, + "timing_ms": {"preprocess": 3.8, "device": 105.3, "postprocess": 3.2, "total": 122.8, "decode": 8.5, "model_call": 112.5}, + "num_poses": 80, + "columns": ["x", "y", "yaw", "cos", "sin", "velocity", "acceleration"], + "trajectory": [[0.365, 0.01, -0.017, 0.9962, -0.0169, 3.8059, -0.0431], [0.7631, -0.0022, -0.0414, 0.9966, -0.0413, 3.8016, -0.5485], [1.1524, -0.0282, -0.0675, 0.9942, -0.0674, 3.7468, -0.4708], "... 77 more rows"], + "turn_indicator": {"command": 1, "command_name": "DISABLE", "keep_selected": true, "held": false, "logits": [-15.960084915161133, -5.125061511993408, -5.230083465576172, -0.7658059597015381, 5.504596710205078], "probabilities": [1.651601411190029e-09, 8.38487030705437e-05, 7.548934809165075e-05, 0.006556871347129345, 0.993184506893158]}, + "predicted_agents": {"format": "npz", "key": "predicted_agents", "dtype": "float32", "shape": [88, 80, 5], "data": ""} +} +``` + +- `trajectory`: 80 points at 0.1-8.0 s in `base_link` (the ego frame of the input tensors), post-processed like the + node's `~/output/trajectory`: velocity from consecutive 3-D points / 0.1 s, forward moving average over + `velocity_smoothing_window` points, force stop (poses frozen once the smoothed speed falls below `stopping_threshold` + while the ego moves), acceleration by finite difference. `yaw` is what `tf2::getYaw` reads from the node's + (unnormalised) quaternion of the raw network `cos` / `sin`, which are returned as they are. +- `turn_indicator`: `command` 0 NO_COMMAND, 1 DISABLE, 2 ENABLE_LEFT, 3 ENABLE_RIGHT (when the head selects KEEP, the + command repeats the last input report `turn_indicators[0, 30]`); the 5 raw `logits` (NONE, DISABLE, LEFT, RIGHT, KEEP) + and the decision's `probabilities`. No hold window is applied (see 3.5). +- `predicted_agents`: one 80-point path (x, y, yaw, cos, sin) per non-empty neighbour row, in input order + (`meta.predicted_agent_rows` gives the rows). +- `meta`: `force_stop`, `time_from_start_s`, `valid_counts` (entities the encoder saw), and with + `return_denoising_steps` the ego row of the 11 solver iterates (`denoising_steps`, `[11, 81, 4]`, the node's + `~/debug/denoising_steps`). + +`timing_ms`: `decode` (base64 + npz parsing + the schema check), `preprocess` (the node's normalization and the encoder's +host features), `device` (host tensors + H2D + one trace replay + D2H), `postprocess`, `model_call`, `total` (server +side, after the body arrived). + +### 3.3 Errors + +**400** undecodable or malformed input (bad base64, a missing or unknown tensor, a wrong shape, a non-finite value, +unknown or out-of-range `params`), **422** schema violation (unknown field, wrong type), **503** while starting (or +when the boot failed), **500** `inference failed: : ` if the device call raises. + +### 3.4 Environment the app reads (lifespan only, never at import) + +| var | set by | meaning / default | +|---|---|---| +| `HF_MODEL` | launcher (`weights.repo`) | weights repo id; default `AutowareFoundation/diffusion_planner` | +| `TT_MODEL_WEIGHTS_REVISION` | launcher (`weights.revision`) | pinned commit; `TT_WEIGHTS_REVISION` (`serve.env`) is the same for older tt-model | +| `DIFFUSION_PLANNER_WEIGHTS_DIR` | you (offline) | local weights directory; overrides the Hub | +| `TT_MESH_SHAPE` | launcher (`runtime.mesh_shape_env`) | `1x1`; anything else -> RuntimeError at startup | +| `TT_DEVICE_ID` | you | chip to open, default 0 | +| `DIFFUSION_PLANNER_DISPATCH` | `serve.env` | `eth` (default) \| `worker` (A/B only, 11×10) \| `auto` (ETH if the patch is present) | +| `DIFFUSION_PLANNER_NUM_CQS` | `serve.env` | `1` (2 CQs replay ~10 % slower on ETH, OPT_BASELINE.md) | +| `DIFFUSION_PLANNER_VARIANT` | `serve.env` | `default` (the only v5.0 graph) | +| `DIFFUSION_PLANNER_LN_FP32`, `_SPLIT_MATMUL`, `_ATTN_MATMUL`, `_HIDDEN_FP32`, `_ATTN_FP32_ACC` | `serve.env` | the numerics knobs, pinned at their gated defaults (`enc.mixer.*,dec.*`; `enc.island.*,enc.pre.*,dec.*`; `enc.fusion.attn,dec.*`; empty; empty: `tt/config.py` `KNOBS`); changing one changes the numerics and needs the gates re-run | +| `DIFFUSION_PLANNER_PRECISION` | `serve.env` | empty: the precision policy `DEFAULT_PRECISION` of `tt/config.py` (HiFi4 + fp32 accumulation everywhere); extra rules such as `dec.*=HiFi2+fp32` are for experiments | +| `DIFFUSION_PLANNER_WARMUP` | you | JSON list of warm-up variants, `default` or `none` | +| `DIFFUSION_PLANNER_TRACE_REGION`, `DIFFUSION_PLANNER_L1_SMALL`, `DIFFUSION_PLANNER_WORKER_L1_SIZE` | you | device-open overrides (validated values: `DEVICE_DEFAULTS` in `device.py`) | +| `DIFFUSION_PLANNER_MAX_BODY_MB` | you | request size guard, default 256 (the shipped sample's request is 0.15 MB of base64 npz) | +| `TT_METAL_VISIBLE_DEVICES`, `MESH_DEVICE`, `HF_HUB_DISABLE_IMPLICIT_TOKEN` | `serve.env` / launcher | `0`, `P150`, `1` (the weights are public: no token is ever sent) | + +Server-side pipeline: JSON -> `tt_diffusion_planner.io` decoders (schema check) -> `DiffusionPlanner.__call__` (the node's +host pre-processing ported from `planning/autoware_diffusion_planner` -> H2D into 19 persistent device inputs -> +`execute_trace` of the whole plan -> one D2H -> the node's host post-processing) under one lock -> `Output.to_dict()`. + +### 3.5 What the client keeps (the API is stateless) + +Every request is one independent plan. The node keeps state between plans; a client that wants its behaviour over a +sequence of plans keeps the same state and puts it into the request or applies it to the response: + +- **The input tensors.** Converting ROS messages and the Lanelet2 map into the 15 tensors is the client's: the per-UUID + agent buffers with their 0.1 s resampling, the ego history, lane / route / polygon / line-string selection and + encoding, traffic-light states, speed limits, the goal and the turn-indicator report history. +- **The turn-indicator hold window.** The node holds its last non-KEEP command for `turn_indicator_hold_duration` + (1.0 s in its YAML). The server evaluates every request with a fresh manager (no held command). To reproduce the node, + keep the last non-KEEP command and its time, and return it instead while less than 1.0 s has passed; the node's + manager ships in the package (`tt_diffusion_planner.host.postprocess.TurnIndicatorManager`, fed with + `turn_indicator.logits`, the plan time and the last report). +- **The initial solver state `x_T` (`sampled_trajectories`, normalised space).** Zeros is the node's default + (`temperature: [0.0]`). With a temperature > 0 the node draws N(0, 1) x temperature for every element (a fresh + `std::random_device` seed per call); send that noise in `sampled_trajectories`. With the RTC prefix (`delay_step` > 0, + at most 40) the node also writes its previous plan into the ego row, slots t = 0 .. delay_step: the previous poses + re-expressed in the current ego frame, `x` as (x - 10) / 20, `y` as y / 20, `cos` / `sin` as they are. Non-zero + `x_T` was checked on the device against ONNX Runtime (noise of scale 0.5 and 1.0, an RTC prefix: within the gates, + `VERIFICATION_2026-10-08.md`). +- **`delay`** is accepted and ignored, as in the node's multi-step mode. +- **The map frame.** The output is in `base_link` (the ego frame of the inputs); the node transforms it to `map` with the + current ego pose before publishing. + +## 4. Caveats + +- Batch 1, one plan per request; concurrent clients queue on the lock. +- Fixed shapes of the v5.0 export (320 neighbours, 140 lanes, 25 route lanes, 10 polygons, 60 line strings, 31 history + and 80 future steps) and 10 DPM-Solver steps (11 decoder evaluations) are compiled into the trace; every plan computes + the full capacities whatever the scene holds, so the device time does not depend on the scene. +- The node's guidance services (start / stop / centerline guidance) are off (the node's default too); a guided mode + would need the host in the solver loop. +- Warm-up captures the trace inside the lifespan, so READY means warm; the first cold boot pays the ttnn JIT. +- Weights are pinned by sha in `weights.revision`, exported by the launcher and repeated in `serve.env`; + `snapshot_download(..., revision=)` is a cache hit after `serve`'s pre-download and falls back to + `local_files_only=True` if the Hub is unreachable. The sha256 of the four files and `major_version == 5` are checked + at load. +- The runtime image has no host C/C++ compiler: device kernels JIT-compile (sfpi ships in the image); the host code is + numpy. +- `tt-model curl` and the ready card's `/v1/models` hint are OpenAI-shaped and are not this API; use the routes above. diff --git a/VERIFICATION_2026-10-08.md b/VERIFICATION_2026-10-08.md new file mode 100644 index 0000000000000000000000000000000000000000..282c6d4eb2e8ce7ff04cab02469b0ea8877c63c0 --- /dev/null +++ b/VERIFICATION_2026-10-08.md @@ -0,0 +1,210 @@ +# diffusion-planner-p150: independent verification, 2026-10-08 + +## Verdict + +**PASS.** This is the light verification pass of the first publish (`research/PLAN.md` §6.3): the frozen gates re-run, +the baseline numbers re-checked, and every device-open path audited for ETH dispatch. It was made by the agent that +wrote the release docs, which is neither the porter nor the baseline agent, on commit `c0d84f9` (the `code/` of this +repo: ttaw 0.20.0 vendored, the numerics knobs pinned in `serve.env`). Against the baseline commit `5541833`: + +- **Accuracy:** the device suite passes (44 passed; again under alloc tracking), every gate value and all 99 per-scene end-to-end numbers identical to the baseline run and to the port's final run: worst ego 0.313 / 0.143 m (max / mean; gates 1.0 / 0.3 m), turn command 99 / 99, neighbours ≤ 0.086 m (gate 1.5 m). +- **Independent oracle:** against the research pipeline's ONNX Runtime goldens, which the bundle's reference did not produce, all 92 / 92 nuScenes instants and 33 / 33 instants of the scene-0061 sequence pass the same gates (worst ego max 0.313 m). +- **Speed:** the stage bench reproduces `OPT_BASELINE.md`: back-to-back replay 102.03 vs 102.04 ms, one plan 102.10 vs 102.13 ms; the host-bound end to end 115.6 vs 117.9 ms p50. +- **Audit:** ETH dispatch, 1 CQ and the 12×10 grid on every device-open path; no fallback in any run. +- **No numerics change** since the port: the diff `5541833..c0d84f9` is the ttaw re-vendor (0.17.1 -> 0.19.0 -> 0.20.0: on this + bundle's path only `trace.py`'s `pack_outputs`, whose single-row layout for the planner's 117,664-element readback is + unchanged), the `serve.env` pins of the numerics knobs (equal to the code defaults), two host-test changes, the + quickstart picture and documentation. No gate threshold moved. + +The full adversarial verification of the port is `VERIFY_PORT.md` of the porting workspace (round 1, 2026-10-08, PASS +with 9 low findings; not part of this repository): it re-ran every suite from a clean shell, audited the request path, +the gates, the oracle chain and the Autoware fidelity line by line, compared the device with ONNX Runtime goldens that +the bundle's own reference did not produce, and ran five adversarial scenes. Its results that this card relies on are +quoted below, and its findings are resolved as listed in "Review findings". + +## Setup + +| item | value | +|---|---| +| Hardware | one Blackhole p150b, ETH dispatch, 1 CQ, 12×10 compute grid (120 cores), as printed by every run (`dispatch eth`, `fallback null`, `eth_patch true`) | +| tt-metal | `44d66500520` with `patches/tt-metal-eth-dispatch.patch` (sha256 `08d0ddf6…45cc`; the tree's `git diff` against `44d6650` is byte-identical to the patch: 4 files) | +| Weights | `AutowareFoundation/diffusion_planner` @ `423efde67f5414734da43a7ad856c17ceb8b51aa` (tag `v5.0`): the Python API, the quickstart and the server read them from the HF cache at the pinned revision (downloaded without a token, sha256 equal to the workspace copy, checked again at every load); the device tests from the workspace copy | +| Verified commit | `c0d84f9` (ttaw 0.20.0, `common` `89dec49`, vendored from the committed tree; `vendor.py --check`: 0 differences) | +| Baseline commit | `5541833` (`OPT_BASELINE.md`) | +| Workload | `code/tt_diffusion_planner/samples/kashiwanoha_dense.npz` (+ `straight_road.npz` and nuScenes scene-0103_kf14 for the stage bench), batch 1, warm; 99 scenes for the accuracy gates and 125 nuScenes instants for the oracle check (local only) | +| Bench method | `code/scripts/bench.py --iters 100 --warmup 5 --b2b-iters 50 --b2b-rounds 3`, one run, the three scenes of the baseline | +| Served method | uvicorn on the host with every `serve.env` pin of `tt-model.yaml` (the container's app and environment), 5 warm-up + 50 requests of the shipped sample from a loopback client | +| Accuracy reference | the fp32 CPU reference (`code/tt_diffusion_planner/reference/`), itself equal to ONNX Runtime on the shipped ONNX (48 tests, PORT_LOG 5.1); independently, the research pipeline's ONNX Runtime goldens | +| Host | AMD EPYC VM, 8 vCPUs, shared with other agents' jobs (load average 5.42, 7.45 during the runs); host-side stages move with it | +| Logs | `logs/diffusion-planner/docs/` of the workspace (job scripts in `scripts/`) | + +## Measured baseline vs re-check + +| measurement (shipped sample `kashiwanoha_dense` unless noted) | baseline `5541833` (`OPT_BASELINE.md`) | re-check `c0d84f9` | change | +|---|---:|---:|---:| +| back-to-back trace replays, per plan (p50) | 102.04 ms | 102.03 ms | -0.01 ms | +| device trace, one blocking plan (p50) | 102.13 ms | 102.10 ms | -0.04 ms | +| device trace, one blocking plan (min) | 102.05 ms | 102.03 ms | -0.02 ms | +| e2e `model(inputs=arrays)` p50 | 117.90 ms | 115.61 ms | -1.9 % | +| e2e `model(inputs=arrays)` p99 | 134.56 ms | 121.79 ms | -9.5 % | +| e2e `model(inputs=<.npz path>)` p50 | 124.76 ms | 118.77 ms | -4.8 % | +| host pre-processing | 4.75 ms | 3.91 ms | -17.8 % | +| pack (19 persistent inputs) | 0.45 ms | 0.40 ms | -11.8 % | +| host tensors (TILE) | 2.41 ms | 2.05 ms | -14.9 % | +| H2D | 1.14 ms | 0.94 ms | -17.1 % | +| D2H (one packed read) | 0.58 ms | 0.37 ms | -36.3 % | +| host post-processing | 4.32 ms | 3.60 ms | -16.7 % | +| back-to-back replays, `straight_road` | 102.04 ms | 102.04 ms | -0.00 ms | +| back-to-back replays, nuScenes scene-0103_kf14 | 102.04 ms | 102.04 ms | -0.00 ms | +| first call after `from_pretrained` | 118.20 ms | 114.79 ms | -2.9 % | +| `from_pretrained` load (warm JIT cache) | 11.64 s | 10.38 s | -10.8 % | + +AICLK during the re-check: median 1350 MHz (min 1343, max 1350; 760 sysfs samples, sampled during the timed loops); the device rows reproduce the baseline to the 0.05 ms the baseline quotes, the host rows move with the load of the shared host. + +Not in the baseline, measured for the card on `c0d84f9`: + +| measurement | value | +|---|---:| +| `from_pretrained` load with an EMPTY JIT cache (firmware + every kernel compiled); first / second call | 314.6 s; 121 / 119 ms | +| `from_pretrained` load with that cache warm (new process); first / second call | 8.6 s; 120 / 120 ms | +| served `/predict` `timing_ms`, median of 50: decode · preprocess · device · postprocess · total | 8.6 · 3.8 · 105.3 · 3.3 · **123.1 ms** (min total 122.5) | +| served client round trip, loopback (median / min; request 0.15 MB) | 127.5 / 127.0 ms | +| server boot to ready (uvicorn, warm JIT cache) | 10.0 s | +| the bundle's `server/smoke_test.py` against it (the container smoke's check) | `PASS diffusion-planner-p150: profile=- variant=default dispatch=eth grid=12x10 cqs=1 n=80 turn=DISABLE reference=kashiwanoha_dense.reference.json (trajectory ade=0.0109 fde=0.0173; predicted_agents max_abs=6.283) device_ms=109.402 total_ms=133.384 rtt_ms=139.0` | +| served body vs the stored CPU reference (the smoke gates) | PASS ({"trajectory": {"n": 80, "ade": 0.01089988605897666, "fde": 0.017330320250936008}, "predicted_agents": {"max_abs_err": 6.282742738723755, "n": 35200}}); `predicted_agents` is reported, not gated: its largest difference is the yaw column at ±π (-3.1413 vs +3.1415 rad, the same heading) | +| the README's request recipe (`client.py --out req.json` + `curl`) | HTTP body with 80 poses, turn DISABLE; same trajectory as the client: True, same predicted agents: True | +| `examples/quickstart.py` as a user runs it (weights from the HF cache, offline, no token) | rc 0; turn DISABLE; `quickstart.json` and `quickstart_bev.png` written | + +## Accuracy gate results + +`test_pcc_device.py` + `test_e2e_device.py` on `c0d84f9`, one devrun job, `TTAW_GATES_READONLY=1`: 44 passed in 42.29s. +Every value equals the baseline run (`OPT_BASELINE.md`, ttaw 0.17.1) and the port's final run (PORT_LOG job 19, ttaw +0.11.0) to the printed precision, and all 99 per-scene end-to-end numbers are identical to both +(`logs/diffusion-planner/docs/compare_runs.json`). + +| gate (PLAN §2.12; `tests/*.gates.json`) | threshold | value | +|---|---:|---:| +| `enc.ego` PCC vs the fp32 reference (valid rows, 2 samples) | ≥ 0.999 | 0.999972 | +| `enc.neighbor` PCC vs the fp32 reference (valid rows, 2 samples) | ≥ 0.999 | 0.999952 | +| `enc.lane` PCC vs the fp32 reference (valid rows, 2 samples) | ≥ 0.999 | 0.999992 | +| `enc.route` PCC vs the fp32 reference (valid rows, 2 samples) | ≥ 0.999 | 0.999998 | +| `enc.polygon` PCC vs the fp32 reference (valid rows, 2 samples) | ≥ 0.999 | 0.999994 | +| `enc.line_string` PCC vs the fp32 reference (valid rows, 2 samples) | ≥ 0.999 | 0.999998 | +| `enc.goal` PCC vs the fp32 reference (valid rows, 2 samples) | ≥ 0.999 | 0.999991 | +| `enc.ego_shape` PCC vs the fp32 reference (valid rows, 2 samples) | ≥ 0.999 | 1 | +| `enc.turn` PCC vs the fp32 reference (valid rows, 2 samples) | ≥ 0.999 | 0.999995 | +| `enc.encoding` PCC (valid tokens) | ≥ 0.999 | 0.999978 | +| `dec.eval` PCC (teacher-forced decoder, min over 11 evaluations x 2 samples) | ≥ 0.999 | 1 (≥ 0.9999995) | +| `ego.max_err_m`, max over 99 scenes | ≤ 1.0 m | 0.313 m (nuscenes/scene-0103_kf14) | +| `ego.mean_err_m`, max over scenes | ≤ 0.3 m | 0.143 m (nuscenes/scene-0103_kf14) | +| `turn.command_agreement` | = 1.0 | 1.000 (99 / 99) | +| `neighbors.median_max_err_m`, max over scenes | ≤ 1.5 m | 0.086 m (kashiwanoha_dense) | +| diagnostics, .kashiwanoha_dense: `final_x0` valid-agent PCC; turn logits max abs | reported | 0.999994; 0.0398 | +| diagnostics, straight_road: `final_x0` valid-agent PCC; turn logits max abs | reported | 1.000000; 0.0372 | +| replay == eager (bit for bit); prefix constraint exact; API == `/predict` | equal | equal | + +Distribution over the 99 scenes: ego mean displacement median 1.3 cm, 95th percentile 6.2 cm. Under `TT_METAL_TRACE_ALLOC_TRACKING=1` (no program compiled after the first capture): 44 passed in 40.50s. + +Also re-run, CPU only: the host suite on the fake ttnn, 107 passed and 45 skipped (the 44 device tests and the ONNX +Runtime module, which runs in the research venv), after the re-vendor and the test changes of `d521981` (0.20.0 re-vendored in `c0d84f9`). + +### Against an independent oracle + +The p150 outputs of the 92 nuScenes v1.0-mini planning instants and of the 33-plan scene-0061 sequence (2 Hz), produced +through the public API (`model(inputs=...)`), compared with the research pipeline's ONNX Runtime goldens +(`research/diffusion-planner/public_data/goldens/`: ORT 1.30 on the shipped ONNX files with an independent host +normalization and solver port), with the PLAN §2.12 gates on the raw denormalised x0 +(`research/diffusion-planner/public_data/scripts/dp_public_metrics.py gates`): + +| set | instants | all four gates pass | worst ego max | worst ego mean | worst neighbour median-max | turn command identical | logits max abs | +|---|---:|---:|---:|---:|---:|---:|---:| +| nuScenes v1.0-mini planning instants (1 Hz, all 10 scenes) | 92 | 92 / 92 | 0.313 m (scene-0103_kf14) | 0.143 m (scene-0103_kf14) | 0.082 m | 92 / 92 | 0.095 | +| scene-0061 sequence (2 Hz, the demo GIF) | 33 | 33 / 33 | 0.118 m (scene-0061_kf36) | 0.049 m (scene-0061_kf09) | 0.027 m | 33 / 33 | 0.074 | + +`VERIFY_PORT.md` (round 1) also ran five adversarial scenes on the device against ONNX Runtime, which the stored goldens +do not cover (every golden has `x_T = 0` and at most 88 agents): full capacity (320 neighbours, 140 lanes, 25 route +lanes, 10 polygons): ego 0.042 / 0.025 m (max / mean); ego only (every neighbour removed): 0.062 / 0.031 m; `x_T` = 0.5 +N(0, 1): 0.022 / 0.007 m; `x_T` = 1.0 N(0, 1): 0.067 / 0.031 m; an RTC prefix in the ego row of `x_T`: 0.024 / 0.012 +m; turn command equal and outputs finite on all five. + +### Open-loop sanity metrics (not a planning metric) + +The p150 plans of the 92 nuScenes instants against the logged drive, with the research metric script +(`dp_public_metrics.py eval`), next to the CPU reference's numbers (`public_data/metrics/open_loop_cpu_fp32.json`). The +model never saw nuScenes and the inputs have known gaps (no traffic-light states, no speed limits), so this is a sanity +check against one human driver, not a paper metric: + +| group (instants) | ADE@3s p150 / CPU | FDE@3s | ADE@8s | FDE@8s | constant velocity ADE@8s / FDE@8s | plan / logged length (m), p150 / CPU | +|---|---:|---:|---:|---:|---:|---:| +| all (92) | 1.12 / 1.12 | 2.60 / 2.59 | 4.75 / 4.74 | 11.87 / 11.87 | 4.99 / 13.06 | 37.5 / 37.5 vs 44.2 | +| moving (v0 >= 1 m/s) (69) | 1.42 / 1.41 | 3.27 / 3.26 | 5.88 / 5.88 | 14.50 / 14.51 | 6.56 / 17.21 | 48.5 / 48.5 vs 58.8 | +| mini_val (18) | 1.44 / 1.44 | 3.42 / 3.42 | 6.42 / 6.43 | 15.86 / 15.89 | 9.47 / 24.11 | 35.6 / 35.6 vs 38.0 | +| mini_train (74) | 1.04 / 1.04 | 2.40 / 2.39 | 4.34 / 4.33 | 10.90 / 10.89 | 3.90 / 10.37 | 38.0 / 38.0 vs 45.7 | + +Turn-indicator command equal to the logged (debounced) blinker: p150 88 / 92, CPU 88 / 92. Moving plans (v0 ≥ 1 m/s) are 17 % shorter than the logged drive. + +## Disclosed numerics changes + +- None since the port. The configuration is decision 11 of `PORT_LOG.md`: fp32 residual streams and solver state; + HiFi4 + fp32 accumulation for every matmul; the ego / neighbour pre-projection as a pad-relative fp32 island; split + hi / lo matmuls (`DIFFUSION_PLANNER_SPLIT_MATMUL=enc.island.*,enc.pre.*,dec.*`); the fp32 LayerNorm decomposition + (`_LN_FP32=enc.mixer.*,dec.*`); fp32 matmul attention (`_ATTN_MATMUL=enc.fusion.attn,dec.*`); the turn head in fp32; + the other weights and the mixer / fusion hidden activations in bf16. This is the precision policy of PLAN.md §9.3 + (HiFi4, fp32 accumulation, two-term or fp32 weights where the gates need them). +- What the configuration costs (timing only, `OPT_BASELINE.md`): 33.4 ms per plan over the first device round's + defaults, which fail `ego.mean_err_m` (0.347 m on nuScenes scene-0103_kf14, PORT_LOG job 11); 58.3 ms over the + fastest graph, which fails `enc.ego`. The card names this cost as the first optimization target. +- The knobs are now pinned in `serve.env` at these values, and a host test (`test_serve_env_pins_the_numerics`) checks + the pins against `KNOBS.serve_env()`; the served `/info` reported the same options (null). + +## Review findings + +- **No gate loosened.** `tests/test_pcc_device.gates.json` and `tests/test_e2e_device.gates.json` are unchanged since + `3223604` (the port's job 18 freeze); the thresholds equal PLAN §2.12; `GateRegistry` refuses a looser declaration and + the runs used `TTAW_GATES_READONLY=1`. +- **No cached or replayed outputs.** The gates compare trace replays with the fp32 CPU reference's goldens, the oracle + check with ONNX Runtime goldens written by another pipeline; the only TT-vs-TT check is replay == eager (a trace + correctness check, not an accuracy gate). +- **The diff `5541833..c0d84f9`** touches no file of `tt_diffusion_planner/tt/`, `host/`, `reference/`, `api.py`, + `device.py`, `io.py` or `server/`: the ttaw re-vendor, `tt-model.yaml` (`serve.env` pins), `examples/quickstart.py` + and two host tests (`test_api_host.py`: the dead `test_build_says_what_is_missing` replaced by + `test_build_refuses_unknown_compile_params`; `test_bundle_host.py`: `test_serve_env_pins_the_numerics`). The docs + commit on top changes only documentation, media and `tt-model.yaml` comments, card text and one `verify:` line. +- **No hang, timeout, reset or FAULT marker** in the device jobs of this pass. + +Status of the `VERIFY_PORT.md` findings: + +| # | finding | status | +|---|---|---| +| L1 | vendored ttaw outdated (0.11.0 vs common 0.15.1) | resolved: re-vendored 0.17.1 at the baseline step (`5541833`) then 0.19.0 (`d521981`) and 0.20.0 (`c0d84f9`) in this pass; the device suite after the first and the last gave identical gate values and per-scene numbers | +| L2 | numerics knobs not pinned in `serve.env`, no host test | resolved (`d521981`): the five knobs and an empty `DIFFUSION_PLANNER_PRECISION` pinned; `test_serve_env_pins_the_numerics` | +| L3 | publish documents outstanding; template text wrong for a planner | resolved: README, SERVING, PYTHON, OPT_BASELINE, OPT_REPORT, this file, the patch copy and the demo media; the card's risks rewritten; packaging, the container smoke, `build_info.json` and the push remain for the packaging step | +| L4 | dead test `test_build_says_what_is_missing` | resolved (`d521981`): replaced by a check that `_build` refuses unknown compile parameters | +| L5 | "6,408 programs" was a fake-ttnn op count | resolved: `OPT_BASELINE.md` quotes 6,282 programs per plan from the device profiler | +| L6 | op-count evidence only in a session scratchpad | resolved: `logs/diffusion-planner/opcount_fake_r2.log` | +| L7 | stateless API: hold window and `x_T` handling left to the client | resolved by documentation: README "Quickstart", `code/PYTHON.md` "What the caller keeps" (with the node's `TurnIndicatorManager` applied across calls), `SERVING.md` 3.5 | +| L8 | the bf16 SDPA path (`DIFFUSION_PLANNER_ATTN_MATMUL=none`) has no gate run since decision 11 | unchanged: it is not a shipped option (`serve.env` pins the matmul attention); PORT_LOG job 13 shows it fails the mean gate on scene-0103_kf14 together with the decoder split (0.74 / 0.33 m), and `OPT_BASELINE.md` times it only | +| L9 | C20 PCC range quoted as 0.99966-0.99983 | resolved in `PORT_LOG.md` (0.99961-0.99983 over all suite shapes, 0.99972-0.99983 at the planner's shapes); `ttaw/API.md` section 17 (common) keeps the old text until the next functional C20 change, as PORT_LOG section 9 says; the gate (0.9995) is unaffected | + +## p150 ETH-dispatch compliance + +| path | dispatch / CQs / grid | +|---|---| +| Python API `DiffusionPlanner.from_pretrained()` (`ttaw.api_base` → `tt_diffusion_planner.device.open_device`) | **ETH / 1 / 12×10** by default (`DeviceConfig.dispatch = "eth"`, `DEVICE_DEFAULTS["num_command_queues"] = 1`; `DIFFUSION_PLANNER_DISPATCH` / `_NUM_CQS` override). A failing ETH open falls back to WORKER with a `RuntimeWarning`, and `model.info["device"]["fallback"]` names it. Measured in this pass (quickstart, load times, demo runs): `dispatch eth`, `grid 12x10`, `fallback null` | +| HTTP server (`tt_diffusion_planner.server.app`, `config_from_env`) | ETH unless `DIFFUSION_PLANNER_DISPATCH` says otherwise; `/health` and `/info` report the device actually opened. Measured (served bench): `/info` `dispatch eth`, `grid 12x10`, `num_command_queues 1`, `fallback None` | +| `tt-model.yaml` serve env (one serve profile, the default) | `DIFFUSION_PLANNER_DISPATCH=eth`, `_NUM_CQS=1`, `_VARIANT=default` and the numerics pins; `check_bundle.py` refuses a profile whose `*_DISPATCH` is not `eth`; host tests check the pins against `device.py` and `tt/config.py` | +| image `verify:` | asserts the ETH-patch marker `single_chip_arch_1cq_no_dispatch_s` in `/opt/tt-metal/tt_metal/impl/dispatch/topology.cpp` and that ttaw sees the patch (`eth_dispatch_patch_present`) | +| container smoke (the one serve profile) | `server/smoke_test.py` asserts `/info` = `eth` / `12x10` (constants, not options) and the profile's pins (dispatch, CQs, variant, weights revision), and compares the served plan with the stored CPU reference; run against uvicorn on the host in this pass (`PASS diffusion-planner-p150: profile=- variant=default dispatch=eth grid=12x10 cqs=1 n=80 turn=DISABLE reference=kashiwanoha_dense.reference.json (trajectory ade=0.0109 fde=0.0173; predicted_agents max_abs=6.283) device_ms=109.402 total_ms=133.384 rtt_ms=139.0`); to be run on the built image by the packaging step | +| gate tests (`code/conftest.py`) | `open_device(..., allow_fallback=False)`: an ETH open that fails is an error, never a silent WORKER run | +| `code/scripts/bench.py`, `profile_ops.py` | `from_pretrained(dispatch=--dispatch or DIFFUSION_PLANNER_DISPATCH or eth)`; the JSON records the device actually opened (`fallback null` in every run); the WORKER rows of `OPT_BASELINE.md` were explicit A/B runs | +| `code/scripts/bringup_device.py`, `precision_exp.py`, `split_error.py` | `open_device(allow_fallback=False)` (ETH) | +| `examples/quickstart.py` | `from_pretrained(device_id=...)`: ETH / 1 / 12×10 (run in this pass) | +| vendored `ttaw` | 0.20.0 @ `89dec49`, `source_dirty: false`; `common/tools/vendor.py --check`: 0 differences; `check_bundle.py`: no drift warning | +| hard-coded grids | none: `tt/` sets no core grid or program config; the only grid consumer, C20's SDPA program config, reads `compute_with_storage_grid_size()` and is not on the default path (matmul attention) | + +## Not covered by this pass + +The container build, its smoke test on the built image, `build_info.json`, the HF pre-flight and the push (the +packaging step); the Python install through `pip install -e .` (the quickstart ran with `code/` on `PYTHONPATH`, the +same import path as the device tests); the dispatch / CQ matrix and the device profile (taken once, in +`OPT_BASELINE.md`, on the same device code). diff --git a/build_info.json b/build_info.json new file mode 100644 index 0000000000000000000000000000000000000000..373f4aa354a7049cd1606ac01de1331ce296eb16 --- /dev/null +++ b/build_info.json @@ -0,0 +1,48 @@ +{ + "schema": "ttaw-build-info/1", + "bundle": "diffusion-planner-p150", + "recorded_at": "2026-10-09T04:48:03Z", + "image": { + "tag": "tt-model/diffusion-planner-p150:3b96d8ea7190", + "digest": "sha256:3b96d8ea71902fe6a00f1792dd41290839b7758070431464cd613cde3f6bf909", + "built_at": "2026-10-09T04:37:19+00:00", + "code_sha256": "c0e7abb7888098a9319ab5c66a10c4fd4009fc1542326b26507554c04f05a481", + "size_bytes": 4069025450, + "layers": 23 + }, + "base_images": { + "build": { + "stage": "prep", + "ref": "ghcr.io/tenstorrent/tt-metal/tt-metalium/ubuntu-22.04-dev-amd64:latest", + "digest": "sha256:df9d279c7f85c17c6fad982d196802682d669cca1b7ced9cbaad8181339cd5fc", + "source": "docker build log", + "local_tag_matches": false, + "local_id": "sha256:3fd1e6013e658c65df1bc6b084543e7174d854c15af97373e493df0c1787c393", + "created": "2026-10-07T01:17:29.186036377Z", + "note": "the local tag moved after this build; the digest above is the one the image was built from" + }, + "runtime": { + "stage": "runtime", + "ref": "docker.io/library/ubuntu:22.04", + "digest": "sha256:5ec03bb3441e8b0bf3b4f9cd4629a1ae763010dc3035bb8da3ae6cf026486401", + "source": "docker build log", + "local_tag_matches": null, + "note": "not in the local image store (BuildKit pulled it into its cache)" + } + }, + "tt_metal": { + "sha": "44d66500520fda9f2c7060c0f6b41ec48f7ab37e", + "describe": "v0.80.0-dev20261006-78-g44d6650052-dirty", + "dirty": true, + "scm_version": "0.65.2.dev11169+g44d66500520", + "mode": "local", + "remote": "https://github.com/tenstorrent/tt-metal.git", + "branch": "main", + "pushed": true + }, + "tools": { + "tt_model": "0.1.0", + "docker": "29.8.2" + }, + "notes": "tt-model's FROM tags float (build: tt-metalium dev image :latest, runtime: ubuntu:); these are the digests this image was built from. A moved build base costs a cold C++ build." +} diff --git a/code/PYTHON.md b/code/PYTHON.md new file mode 100644 index 0000000000000000000000000000000000000000..0d29e7c92256b84088f1eb229c061c46a9e10d0c --- /dev/null +++ b/code/PYTHON.md @@ -0,0 +1,179 @@ +# Python API: Diffusion Planner v5.0 (Autoware diffusion_planner) on Blackhole + +Use this API from Python code (a pipeline, a notebook, a ROS 2 node wrapper). You do not need the HTTP server: the +API and the server share the decoders, the device trace and the post-processing, so the outputs and the speed are +the same. + +## Install + +Install the package on top of an environment that already has `ttnn` (a tt-metal `python_env` at `44d66500520` +with `patches/tt-metal-eth-dispatch.patch`, or the tt-model container). From the root of the model repository (the +directory that holds `pyproject.toml`, `README.md` and `code/`): + +```bash +pip install -e . # the Python API (numpy<2, pillow, pyyaml, onnx, huggingface_hub) +pip install -e ".[server,test]" # + the HTTP server and the tests +``` + +The pip project is the repository's top-level `pyproject.toml`; it installs the package from +`code/tt_diffusion_planner` (there is no `pyproject.toml` inside `code/`, because the container build copies `code/` +over the tt-metal tree). ttnn and torch come from tt-metal and are not declared. + +The package carries `tt_diffusion_planner.ttaw`, the shared code of the Autoware ports to Blackhole (device open, trace +runner, decoders, model base class, HTTP app), vendored at the version recorded in +`code/tt_diffusion_planner/ttaw/VENDORED.json`. + +| You want to run | Extras | +|---|---| +| the Python API | none | +| the HTTP server (`tt_diffusion_planner.server.app`, see `SERVING.md`) | `server` | +| host tests (no device; device tests are skipped): `TT_VISIBLE_DEVICES=none python -m pytest -q code/tt_diffusion_planner/tests` | `server,test` | +| device tests: `python -m pytest -q -s code/tt_diffusion_planner/tests/test_pcc_device.py code/tt_diffusion_planner/tests/test_e2e_device.py` | `test` | + +## Quickstart + +```python +from tt_diffusion_planner import DiffusionPlanner + +with DiffusionPlanner.from_pretrained(device_id=0) as model: + out = model(inputs="code/tt_diffusion_planner/samples/kashiwanoha_dense.npz") +print(out.to_dict()) # the POST /predict body +``` + +`examples/quickstart.py` runs the same snippet, writes `quickstart.json` and a bird's-eye view of the input and the +plan (`quickstart_bev.png`). + +## `DiffusionPlanner.from_pretrained(...)` + +```python +DiffusionPlanner.from_pretrained( + model_id=None, # HF repo or a local directory with the weights files; default AutowareFoundation/diffusion_planner + *, + revision=None, # default for the default repo: the validated commit 423efde67f5 (tag v5.0) + variant=None, # "default" (the only v5.0 graph); default $DIFFUSION_PLANNER_VARIANT or "default" + device_id=None, # chip to open; default $TT_DEVICE_ID or 0 + device=None, # an already-opened ttnn device (tt_diffusion_planner.device.open_device); close() does not close it + dispatch=None, # "eth" (p150 target, 12x10 grid) | "worker" (A/B only, 11x10) | "auto"; default $DIFFUSION_PLANNER_DISPATCH or "eth" + num_command_queues=None, # default $DIFFUSION_PLANNER_NUM_CQS or 1 + weights_dir=None, # explicit local weights directory; no Hub access + warmup_variants="default", # trace variants to capture now; see "Warm-up" + verbose=False, + precision=None, # the only compile parameter: extra precision-policy rules, e.g. "dec.*=HiFi2+fp32" (experiments only) +) -> DiffusionPlanner +``` + +What it does: resolves the weights first, so a Hub problem never claims the chip (`weights_dir` > +`$DIFFUSION_PLANNER_WEIGHTS_DIR` > a local `model_id` directory > the HF snapshot at the pinned revision, restricted to +the three v5.0 ONNX files and `diffusion_planner.param.json`, with an offline fallback to the cache; the sha256 of every +file and the weights' `major_version == 5` are checked), opens the chip (ETH dispatch, 12×10, 1 CQ; the other open +parameters are `DEVICE_DEFAULTS` in `tt_diffusion_planner/device.py`, overridable with `DIFFUSION_PLANNER_*`), reads the +ONNX initializers as data, uploads the weights and constants (48.1 MB), builds the graph, then compiles and captures +the metal trace. If ETH dispatch cannot open (tt-metal without the patch), it warns and falls back to WORKER dispatch +(`model.info["device"]["fallback"]` names it). Any other keyword argument is a `TypeError`. + +The numerics are not arguments: the published configuration is the default of the `DIFFUSION_PLANNER_LN_FP32`, +`_SPLIT_MATMUL`, `_ATTN_MATMUL`, `_HIDDEN_FP32` and `_ATTN_FP32_ACC` knobs (`tt_diffusion_planner.tt.config.KNOBS`, +pinned in `tt-model.yaml` `serve.env`). Setting one of them in the environment changes the graph and invalidates the +accuracy figures of the card until the gates are re-run. + +## Warm-up + +`from_pretrained` returns a warm model: it builds the graph, runs the plan once eagerly (this first run compiles every kernel into the JIT cache), then captures the whole plan as one metal trace (`warmup_variants="default"`: the variant `plan`) with program-cache misses forbidden, so no later call compiles anything. `model.warmup()` is idempotent; `warmup_variants="none"` defers the capture to `model.warmup()`. + +Measured on the shipped sample (`model.info["warmup_ms"]`, 2026-10-09): the load takes 315 s with an empty JIT cache and 8.6 s with a warm one (build 0.53 s: the ONNX initializers read and 48.1 MB of weights and constants uploaded; warm-up and capture 3.7 s; the rest is the device open). The first call then takes 120 ms and the second 120 ms (the stage bench's steady state: 118 ms p50 for decoded arrays, 125 ms for an `.npz` path). The trace holds 74.6 MB of DRAM (`trace_region_size` 192 MiB). + +## Call: `model(...)` + +| Argument | Type | Description | +|---|---|---| +| `inputs` | mapping / `.npz` path / bytes / JSON envelope | the 15 raw planner tensors (see "Input types") | +| `velocity_smoothing_window` | int, 1..79, default 8 | forward moving average of the trajectory velocity, in points | +| `stopping_threshold` | float >= 0, default 0.3 | force stop below this smoothed speed (m/s), when the ego moves | +| `turn_indicator_keep_offset` | float, default -1.25 | added to the KEEP logit before the turn-indicator decision | +| `return_denoising_steps` | bool, default False | add the ego row of the 11 solver iterates (`out.meta["denoising_steps"]`, `[11, 81, 4]`, the node's `~/debug/denoising_steps`) | + +### Input types + +- `inputs=`: the 15 raw tensors of the Autoware node's `DiffusionPlannerCore::create_input_data()` (batch 1, + float32, ego `base_link` frame, BEFORE normalization; names and shapes in `tt_diffusion_planner.INPUT_SCHEMA`): + a `{name: array}` mapping (numpy or torch), an `.npz` path or its bytes, or the `/predict` envelope + `{"format": "npz", "data": }` / `{"format": "json", "arrays": {...}}`. Names, shapes and finite values are + checked (`InputError`). `tt_diffusion_planner.load_inputs(source)` is the same decoder. +- Any other input (`points`, `images`, `calibration`, ...) is refused (`InputError`). + +### What the caller keeps (the API is stateless) + +One call is one independent plan. The Autoware node keeps state between plans; to reproduce it over a sequence of +plans, the caller keeps the same state (SERVING.md 3.5 has the details): + +- **The tensors.** The node's pre-processing from ROS messages and the Lanelet2 map (per-UUID agent buffers and their + 0.1 s resampling, the ego history, lane / route / polygon / line-string selection and encoding, traffic lights, + speed limits, goal, turn-indicator report history) is not part of the bundle. +- **The turn-indicator hold window** (`turn_indicator_hold_duration`, 1.0 s in the node's YAML). Each call decides + with a fresh manager; apply the node's hold across calls with the node's own manager: + + ```python + from tt_diffusion_planner.host.postprocess import TurnIndicatorManager + + manager = TurnIndicatorManager() # hold 1.0 s, KEEP offset -1.25 (the node's YAML) + out = model(inputs=tensors) + decision = manager.evaluate(out.turn_indicator["logits"], stamp_s=now_s, prev_report=int(tensors["turn_indicators"][0, 30])) + command = decision.command # the held command while less than 1.0 s has passed + ``` + +- **The initial solver state** `sampled_trajectories` (`x_T`, normalised space): zeros is the node's default + (`temperature: [0.0]`); for a temperature > 0 send N(0, 1) x temperature; for the RTC prefix (`delay_step` > 0) put + the previous plan into the ego row, slots t = 0 .. delay_step (x as (x - 10) / 20, y as y / 20, cos / sin as they are, + in the current ego frame). `delay` is accepted and ignored (the node's multi-step mode never reads it). +- **The map frame.** Outputs are in `base_link`; the node transforms them with the current ego pose to `map`. + +## Output + +`model(...)` returns a `tt_diffusion_planner.Output` (= `ttaw.outputs.Trajectory`); `out.to_dict()` is exactly the `POST /predict` body (SERVING.md section 3.2). + +| field | type | meaning | +|---|---|---| +| `poses` | float32 `[80, 7]` | the ego trajectory at 0.1-8.0 s in `base_link`: x, y, yaw, cos, sin, velocity, acceleration (`out.columns`), post-processed like the node's `~/output/trajectory` | +| `turn_indicator` | dict | `command` (0 NO_COMMAND, 1 DISABLE, 2 ENABLE_LEFT, 3 ENABLE_RIGHT), `command_name`, `keep_selected`, `held` (always false: no hold window), the 5 raw `logits` (NONE, DISABLE, LEFT, RIGHT, KEEP), the decision's `probabilities` | +| `predicted_agents` | float32 `[N, 80, 5]` | x, y, yaw, cos, sin of each non-empty neighbour row, in input order | +| `meta` | dict | `predicted_agent_rows` (the rows of `predicted_agents`), `predicted_agent_columns`, `force_stop`, `time_from_start_s`, `valid_counts` (the entities the encoder saw), and with `return_denoising_steps` the encoded `denoising_steps` | +| `timing_ms` | dict | `preprocess`, `device` (host tensors + H2D + replay + D2H), `postprocess`, `total` | + +`out.to_dicts()` gives one `{x, y, yaw, cos, sin, velocity, acceleration}` dict per trajectory point; `out.to_dict("npz")` adds the poses as a lossless base64 NPZ array. + +## Lifetime and information + +- `model.close()` releases the trace and the persistent device tensors and closes the chip if the model opened it; + idempotent. `with` calls it for you; an unclosed model is closed when Python exits. +- `model.info`: weights (repo, tag, revision, path), device (dispatch, grid, CQs, fallback), variant, warm variants, + warm-up times, runtime parameter defaults, the input schema, the numerics options and precision policy in effect, the + trace (variants, persistent inputs, trace buffers in MB). +- Calls from several threads are safe: the device calls are serialised. One model per process per chip. + +## Speed + +Warm calls, batch 1, ETH dispatch, 1 CQ, 12×10, the pinned numerics (`code/scripts/bench.py`, 100 iterations; the numbers of `OPT_BASELINE.md`, 2026-10-08, on a shared host; p50, with p99 in brackets): + +| stage | shipped sample `kashiwanoha_dense` | +|---|---:| +| `.npz` decode + schema check (path inputs only) | 7.02 (18.20) ms | +| host pre-processing (the node's normalization, the encoder's host features, masks, `x_T`) | 4.75 (12.55) ms | +| packing the 19 persistent trace inputs · ttnn host tensors · H2D | 0.45 · 2.41 · 1.14 ms | +| **device trace, one blocking plan** | **102.13** (104.81) ms | +| D2H (one packed read, 460 KB) | 0.58 ms | +| host post-processing (trajectory, predicted paths, turn decision) | 4.32 (11.54) ms | +| **`model(inputs=arrays)` end to end** | **117.90** (134.56) ms | +| `model(inputs=<.npz path>)` | 124.76 (147.74) ms | +| back-to-back replays (device time per plan) | 102.04 ms = 9.80 plans/s | + +The device time does not depend on the scene: every plan computes the full capacities (re-checked on c0d84f9: kashiwanoha_dense 102.10, straight_road 102.11, a nuScenes instant 102.09 ms). Throughput above one plan per ~118 ms needs pipelining of the host work of neighbouring requests (not implemented); 2 CQs do not help a synchronous request (`OPT_BASELINE.md`). Where the time goes and what comes next: `OPT_REPORT.md`. + +## Limits + +- Batch 1 on the chip; one model per process. +- Fixed shapes of the v5.0 export (320 neighbours, 140 lanes, 25 route lanes, 10 polygons, 60 line strings, 31 history + and 80 future steps) and 10 DPM-Solver steps (11 decoder evaluations) are compiled into the trace; every plan computes + the full capacities, so the device time does not depend on the scene. +- The node's guidance services (start / stop / centerline guidance) are not available (the node's default is off). +- Accuracy is agreement with the fp32 CPU reference of the same network (README "Demo & Performances"); the planner's + driving quality is the weights' (trained by TIER IV on data that is not public). diff --git a/code/conftest.py b/code/conftest.py new file mode 100644 index 0000000000000000000000000000000000000000..fb59d40a93d13de0b5592238ff01030a8611fe92 --- /dev/null +++ b/code/conftest.py @@ -0,0 +1,59 @@ +# SPDX-License-Identifier: Apache-2.0 +"""pytest fixtures of diffusion-planner-p150 (no dependency on tt-metal's own conftest). + +- ``device`` (session): one chip opened like the published numbers (``tt_diffusion_planner.device.open_device``: + ETH dispatch, 12x10 grid, the port's validated sizes and CQs). ``--device-id N`` or ``TT_DEVICE_ID`` selects + the chip (default 0); ``DIFFUSION_PLANNER_DISPATCH=worker`` is the A/B opt-in, and the other ``DIFFUSION_PLANNER_*`` + variables apply as for the server. A failing ETH open is an error here, never a silent WORKER fallback: gates + are only valid on the published setup. +- Tests marked ``device`` are skipped when ttnn is missing or ``TT_VISIBLE_DEVICES=none`` (host-only runs). + On the shared workspace box run them through the lock and name the test files: + ``bin/devrun -t 1800 -- python -m pytest -q -s code/tt_diffusion_planner/tests/test_pcc_device.py``. +""" +from __future__ import annotations + +import gc +import importlib.util +import os + +import pytest + + +def pytest_addoption(parser): + parser.addoption("--device-id", action="store", default=None, help="chip id (default $TT_DEVICE_ID or 0)") + + +def pytest_configure(config): + config.addinivalue_line("markers", "device: needs a Tenstorrent chip (skipped on host-only runs)") + + +def _no_device_reason(): + if os.environ.get("TT_VISIBLE_DEVICES", "").lower() == "none": + return "TT_VISIBLE_DEVICES=none (host-only run)" + if importlib.util.find_spec("ttnn") is None: + return "ttnn is not installed" + return None + + +def pytest_collection_modifyitems(config, items): + reason = _no_device_reason() + if reason: + skip = pytest.mark.skip(reason=reason) + for item in items: + if "device" in item.keywords: + item.add_marker(skip) + + +@pytest.fixture(autouse=True) +def _gc_between_tests(): + gc.collect() + + +@pytest.fixture(scope="session") +def device(request): + from tt_diffusion_planner.device import close_device, open_device + + cli = request.config.getoption("--device-id") + dev = open_device(int(cli) if cli is not None else None, allow_fallback=False) + yield dev + close_device(dev) diff --git a/code/models/common/lightweightmodule.py b/code/models/common/lightweightmodule.py new file mode 100644 index 0000000000000000000000000000000000000000..6743db9026f0547737bda649db0b3ab8c78289c0 --- /dev/null +++ b/code/models/common/lightweightmodule.py @@ -0,0 +1,12 @@ +# SPDX-FileCopyrightText: © 2023 Tenstorrent USA, Inc. + +# SPDX-License-Identifier: Apache-2.0 + + +class LightweightModule: + """Torch modules add a surprising amount of host overhead for attribute + access and method calls. This class is a lightweight alternative that + just wraps a forward function for now.""" + + def __call__(self, *args, **kwargs): + return self.forward(*args, **kwargs) diff --git a/code/scripts/README.md b/code/scripts/README.md new file mode 100644 index 0000000000000000000000000000000000000000..fc1e4f2c6a2d9340da5704f3a11108010663b326 --- /dev/null +++ b/code/scripts/README.md @@ -0,0 +1,16 @@ +# code/scripts + +| script | purpose | device | +|---|---|---| +| `bench.py` | stage breakdown of warm plans (load / host_pre / pack / host_in / H2D / trace / D2H / host_post / e2e / b2b, p50 / p99 / min, AICLK) on one or more scene `.npz` files, `--dispatch` / `--num-cqs` for the dispatch / CQ matrix (the card and `OPT_BASELINE.md` numbers) | yes, via `bin/devrun` | +| `profile_ops.py` | one eager plan (stage `m:` and layer-kind `c:` signposts) + one traced replay between signposts under the device profiler (`python -m tracy -r -p -v --op-support-count 16000 ...`: a plan is 6,282 programs, over the default 1,000-program buffer) | yes | +| `bringup_device.py`, `precision_exp.py`, `split_error.py` | the port's numerics experiments (module PCC per configuration, LayerNorm / split-matmul / attention variants, where the plan error of a scene comes from: device encoder vs device decoder); their logs and findings are in `PORT_LOG.md` (workspace) | yes | +| `ref_golden.py` | goldens of the fp32 CPU reference (per-module taps, final outputs, the stored `/predict` references next to the samples) for the device tests; research venv (onnxruntime): `tools/research-venv/bin/python code/scripts/ref_golden.py` | no | +| `container_smoke.sh` | serve ONE serve profile of the built package (`--profile NAME`; default profile otherwise), run `server/smoke_test.py` against it (asserts ETH dispatch, the 12x10 grid and the profile's pins; compares with the stored CPU reference), keep the evidence (container log, `/info`, the `/predict` output and the result, in `logs/smoke/` of the repo or `--log-dir DIR`), always stop it; one `bin/devrun -t 3600 -k 150` window per profile | yes | +| `fetch_samples.sh` | downloads public sample data that may not be redistributed (sha256-checked); diffusion-planner-p150 has none to fetch: both shipped samples are Apache-2.0, and the nuScenes-derived planning instants of the accuracy tables live in the tt-models workspace only | no | + +Everything here ships in `code/` on the Hub and inside the image (`source.extra_code` lists `scripts`). + +The card's demo media (`media/`) were rendered from p150 outputs by workspace scripts that are not shipped, because +most of the data they read (nuScenes) may not be redistributed (`media/ATTRIBUTION.md`). `examples/quickstart.py` +writes a simple render (`quickstart_bev.png`: the input tensors and the plan, bird's-eye view) for any scene `.npz`. diff --git a/code/scripts/bench.py b/code/scripts/bench.py new file mode 100644 index 0000000000000000000000000000000000000000..4827bfcea77b4e677ce135211165236397467f45 --- /dev/null +++ b/code/scripts/bench.py @@ -0,0 +1,207 @@ +#!/usr/bin/env python3 +# SPDX-License-Identifier: Apache-2.0 +"""Stage breakdown of warm plans -- the numbers OPT_BASELINE.md / OPT_REPORT.md / the card quote. + + bin/devrun -t 900 -- python code/scripts/bench.py --iters 100 --json out.json + bin/devrun -t 900 -- python code/scripts/bench.py --dispatch worker --num-cqs 2 --iters 100 --json out.json + bin/devrun -t 900 -- python code/scripts/bench.py --input --input ... + +``--input`` takes the planner tensors as an ``.npz`` with the 15 ``INPUT_SCHEMA`` names (the shipped samples) or with +``raw/`` keys (the research / public-data scene files); default: the shipped ``kashiwanoha_dense.npz``. +Per input, ``ttaw.profiling.StageBench`` collects ``--iters`` warm iterations of each stage (p50 / p99 / mean / min): + +- ``load``: decoding the ``.npz`` + the ``INPUT_SCHEMA`` check (``model(inputs=)`` does it outside timing_ms); +- ``host_pre``: the node's pre-processing (``host.prepare``: normalization, speed masks, encoder host features, + decoder masks, the solver's initial state); +- ``pack``: the 19 persistent trace inputs (``tt.inputs.plan_inputs``); +- ``host_in``: their ttnn host tensors (fp32 / bf16 TILE, ``ttaw.tensors.to_host_tensor``); +- ``h2d``: the upload into the persistent device inputs + device sync; +- ``trace``: one replay of the ``plan`` trace + device sync (the device latency of one plan); +- ``d2h``: the one packed readback (``final_x0`` + logits + the ego rows of the 11 iterates); +- ``host_post``: the node's post-processing (``host.make_output``); +- ``e2e``: ``model(inputs=)``, the in-process API call (schema check included); +- ``e2e_path``: ``model(inputs=)`` (adds ``load``; schema-named ``.npz`` files only); +- ``b2b``: back-to-back replays with no host work in between (device time per plan), ``--b2b-rounds`` rounds of + ``--b2b-iters`` replays; ``plans_per_s`` = 1000 / median b2b. + +Also: ``model(...).timing_ms`` (preprocess / device / postprocess / total), the first call after ``from_pretrained`` +and the load time (weights, build, warm-up + capture), AICLK / power / temperature sampled from sysfs during the +timed loops (``ttaw.profiling.AiclkSampler``), the staged path checked bit for bit against ``model()``, the device +configuration (dispatch, CQs, grid) and the numerics options in effect (``DIFFUSION_PLANNER_*`` knobs). Always quote +the configuration line with the numbers. +""" +from __future__ import annotations + +import argparse +import json +import statistics +import time +from pathlib import Path +from typing import Any, Dict, Optional + +import numpy as np + +from tt_diffusion_planner import DiffusionPlanner + +SAMPLE = Path(__file__).resolve().parents[1] / "tt_diffusion_planner" / "samples" / "kashiwanoha_dense.npz" + + +def load_scene(path: str) -> Dict[str, Any]: + """``{"name", "path", "arrays", "schema_npz"}``: the 15 raw tensors of a schema-named or ``raw/``-prefixed npz.""" + from tt_diffusion_planner.reference import config as C + + p = Path(path) + with np.load(p, allow_pickle=False) as z: + files = set(z.files) + if all(k in files for k in C.INPUT_NAMES): + arrays, schema_npz = {k: np.array(z[k]) for k in C.INPUT_NAMES}, True + elif all(f"raw/{k}" in files for k in C.INPUT_NAMES): + arrays, schema_npz = {k: np.array(z[f"raw/{k}"]) for k in C.INPUT_NAMES}, False + else: + raise SystemExit(f"{p}: neither the INPUT_SCHEMA names nor raw/ keys") + name = p.stem[len("golden_"):] if p.stem.startswith("golden_") else p.stem + if p.parent.name not in ("samples", "ort"): + name = f"{p.parent.name}/{name}" + return {"name": name, "path": str(p), "arrays": arrays, "schema_npz": schema_npz} + + +def counts(arrays: Dict[str, np.ndarray]) -> Dict[str, int]: + """Valid entities of a scene (non-empty rows), for the report.""" + def rows(a, axis): + return int(np.any(np.abs(a) > 0, axis=axis).sum()) + return {"neighbors": rows(arrays["neighbor_agents_past"][0], (1, 2)), "lanes": rows(arrays["lanes"][0], (1, 2)), + "route_lanes": rows(arrays["route_lanes"][0], (1, 2)), "polygons": rows(arrays["polygons"][0], (1, 2)), + "line_strings": rows(arrays["line_strings"][0], (1, 2))} + + +def unpack(out: Dict[str, np.ndarray]) -> Dict[str, Any]: + """The packed readback -> the raw outputs of ``TtDiffusionPlanner.forward`` (same reshapes).""" + from tt_diffusion_planner.reference import config as C + from tt_diffusion_planner.tt import config as T + + final = out["final_x0"].reshape(T.AGENTS, T.STATE_COLS)[:C.MAX_NUM_AGENTS] + steps = out["ego_steps"].reshape(-1, T.STATE_COLS) + return {"final_x0": final.reshape(C.MAX_NUM_AGENTS, C.OUTPUT_T + 1, C.POSE_DIM).astype(np.float32), + "logit": out["logit"].reshape(-1)[:C.TURN_INDICATOR_OUTPUT_DIM].astype(np.float32), + "denoising_steps": [s.reshape(1, C.OUTPUT_T + 1, C.POSE_DIM).astype(np.float32) for s in steps]} + + +def bench_scene(model, scene: Dict[str, Any], a: argparse.Namespace) -> Dict[str, Any]: + import ttnn + + from tt_diffusion_planner.host import pipeline as hp + from tt_diffusion_planner.reference import config as C + from tt_diffusion_planner.tt import inputs as I + from tt_diffusion_planner.ttaw.io import load_named_arrays + from tt_diffusion_planner.ttaw.profiling import AiclkSampler, StageBench, time_b2b + from tt_diffusion_planner.ttaw.tensors import to_host_tensor + + runner, dev = model.runner, model.device + arrays, params = scene["arrays"], model.validate_params({}) + obs = model.normalization.observation + bench = StageBench(f"diffusion-planner {scene['name']}") + timing: Dict[str, list] = {} + sync = lambda: ttnn.synchronize_device(dev) # noqa: E731 + for _ in range(a.warmup): + ref = model(inputs=arrays) + with AiclkSampler(chip=a.chip, interval_s=0.05) as clk: + for _ in range(a.iters): # the in-process API call + with bench.stage("e2e"): + ref = model(inputs=arrays) + for k, v in ref.timing_ms.items(): + timing.setdefault(k, []).append(v) + slots = {k: runner._slot(k, "input") for k in I.INPUT_SPECS} + for _ in range(a.iters): # the same path, stage by stage + if scene["schema_npz"]: + with bench.stage("load"): + raw = load_named_arrays(scene["path"], C.INPUT_SCHEMA) + else: + raw = load_named_arrays(arrays, C.INPUT_SCHEMA) + with bench.stage("host_pre"): + prep = hp.prepare(raw, obs) + with bench.stage("pack"): + packed = I.plan_inputs(prep) + with bench.stage("host_in"): + host = {k: to_host_tensor(v, slots[k].dtype, slots[k].layout, shape=slots[k].shape) + for k, v in packed.items()} + with bench.stage("h2d"): + runner.upload(host) + sync() + with bench.stage("trace"): + runner.replay("plan") + sync() + with bench.stage("d2h"): + out = runner.read("plan") + with bench.stage("host_post"): + res = model._postprocess(unpack(out), prep, params) + if scene["schema_npz"]: + for _ in range(a.iters): + with bench.stage("e2e_path"): + model(inputs=scene["path"]) + rounds = [time_b2b(lambda: runner.replay("plan"), sync, n=a.b2b_iters, warmup=3) + for _ in range(a.b2b_rounds)] + for r in rounds: + bench.add("b2b", r) + same = bool(np.array_equal(res.poses, ref.poses) and np.array_equal(res.predicted_agents, ref.predicted_agents) + and res.turn_indicator["command"] == ref.turn_indicator["command"]) + summary = bench.summary() + b2b = statistics.median(rounds) + return {"name": scene["name"], "path": scene["path"], "valid": counts(arrays), "stages_ms": summary, + "timing_ms": {k: {"p50": float(np.percentile(v, 50)), "p99": float(np.percentile(v, 99)), + "min": float(min(v))} for k, v in timing.items()}, + "b2b_rounds_ms": rounds, "plans_per_s_b2b": 1000.0 / b2b, "aiclk": clk.summary(), + "staged_equals_model": same, "turn_command": int(ref.turn_indicator["command"]), + "table": bench.table()} + + +def main() -> None: + ap = argparse.ArgumentParser(description=__doc__.split("\n")[0]) + ap.add_argument("--iters", type=int, default=100) + ap.add_argument("--warmup", type=int, default=5) + ap.add_argument("--b2b-iters", type=int, default=50) + ap.add_argument("--b2b-rounds", type=int, default=3) + ap.add_argument("--input", action="append", default=None, help="scene .npz (repeatable)") + ap.add_argument("--dispatch", default=None, choices=["eth", "worker"]) + ap.add_argument("--num-cqs", type=int, default=None, choices=[1, 2]) + ap.add_argument("--chip", type=int, default=0, help="sysfs chip index for the AICLK sampler") + ap.add_argument("--tag", default="") + ap.add_argument("--json") + a = ap.parse_args() + from tt_diffusion_planner.reference.weights import find_weights_dir + + scenes = [load_scene(p) for p in (a.input or [str(SAMPLE)])] + wd = find_weights_dir() # None: from_pretrained resolves the pinned HF snapshot + t0 = time.perf_counter() + with DiffusionPlanner.from_pretrained(dispatch=a.dispatch, num_command_queues=a.num_cqs, + weights_dir=str(wd) if wd else None) as model: + load_s = time.perf_counter() - t0 + t1 = time.perf_counter() + model(inputs=scenes[0]["arrays"]) # first call after from_pretrained (traces captured) + first_ms = (time.perf_counter() - t1) * 1e3 + info = model.tt.describe() + res: Dict[str, Any] = { + "tag": a.tag, "config": model.device_info, "iters": a.iters, "load_s": round(load_s, 2), + "warmup_ms": {k: round(v, 1) for k, v in model.warmup_ms.items()}, "first_call_ms": round(first_ms, 2), + "options": info["options"], "precision": info["precision"], "uploaded_mb": info["uploaded_mb"], + "trace_buffers_mb": info["trace_buffers_mb"], + "program_cache_entries": info["trace"].get("program_cache_entries"), + "num_command_queues": info["trace"]["num_command_queues"], "scenes": {}} + for scene in scenes: + r = bench_scene(model, scene, a) + res["scenes"][r["name"]] = r + cfg = model.device_info + print(f"\n## {r['name']} [{cfg.get('dispatch')} {res['num_command_queues']}CQ {cfg.get('grid')}] " + f"valid {r['valid']}") + print(r["table"]) + print("timing_ms p50:", {k: round(v["p50"], 3) for k, v in r["timing_ms"].items()}, + f"| b2b rounds {[round(x, 3) for x in r['b2b_rounds_ms']]} ms -> {r['plans_per_s_b2b']:.2f} plans/s", + f"| aiclk {r['aiclk'].get('aiclk_mhz')}", f"| check: staged == model() {r['staged_equals_model']}", + flush=True) + print(json.dumps({k: v for k, v in res.items() if k != "scenes"}, default=str)) + if a.json: + Path(a.json).parent.mkdir(parents=True, exist_ok=True) + Path(a.json).write_text(json.dumps(res, indent=1, default=str) + "\n") + + +if __name__ == "__main__": + main() diff --git a/code/scripts/bringup_device.py b/code/scripts/bringup_device.py new file mode 100644 index 0000000000000000000000000000000000000000..27a6f0360ffe7a8e2c98705aab77c5c7045361fd --- /dev/null +++ b/code/scripts/bringup_device.py @@ -0,0 +1,167 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Device bring-up of the planner graph (development tool; needs the p150, run through ``bin/devrun``). + + bin/devrun -t 1800 -- python code/scripts/bringup_device.py [--scenes kashiwanoha_dense straight_road] + [--no-capture] [--precision "dec.*=HiFi2+fp32"] [--json logs/diffusion-planner/bringup.json] + +Protocol (PLAN.md 4.4): build ``TtDiffusionPlanner(debug=True)``; per scene run every variant EAGERLY twice +(``encoder_taps``, ``decode_once`` at evaluations 0 / 5 / 10, ``plan``), check the two eager runs are bit-identical and +compare them with the research goldens (``research/diffusion-planner/goldens/.npz``, the fp32 CPU reference: +PCC on valid rows, max abs; final x0 on valid agents; logits); then capture all variants (strict, no program +compiled after the capture) and check replay == eager bit for bit on the same inputs. Prints a table and writes a +JSON report. This is a diagnostic: the frozen gates live in ``tests/test_pcc_device.py`` / ``test_e2e_device.py``. +""" +from __future__ import annotations + +import argparse +import json +import sys +import time +from pathlib import Path + +import numpy as np + +HERE = Path(__file__).resolve() +sys.path.insert(0, str(HERE.parents[1])) + +from tt_diffusion_planner.host import pipeline as hp # noqa: E402 +from tt_diffusion_planner.reference import config as C # noqa: E402 +from tt_diffusion_planner.ttaw.metrics import error_stats # noqa: E402 + +GOLDENS = HERE.parents[4] / "research" / "diffusion-planner" / "goldens" + + +def stats(dev, ref): + s = error_stats(np.asarray(dev, np.float64), np.asarray(ref, np.float64)) + return {"pcc": round(s["pcc"], 7), "max_abs": float(f"{s['max_abs']:.4g}"), "rel_l2": float(f"{s['rel_l2']:.4g}")} + + +def main() -> int: + ap = argparse.ArgumentParser(description=__doc__.split("\n\n")[0]) + ap.add_argument("--scenes", nargs="+", default=["kashiwanoha_dense", "straight_road"]) + ap.add_argument("--no-capture", action="store_true") + ap.add_argument("--precision", default=None) + ap.add_argument("--ln-fp32", default=None, help="module globs (comma-separated) for the fp32 LayerNorm") + ap.add_argument("--hidden-fp32", default=None, help="module globs (comma-separated) for fp32 hidden activations") + ap.add_argument("--split", default=None, help="module globs (comma-separated) for split (hi/lo) fp32 matmuls") + ap.add_argument("--attn-fp32-acc", default=None, help="module globs for SDPA with fp32 accumulation") + ap.add_argument("--attn-matmul", default=None, help="module globs for the fp32 matmul attention") + ap.add_argument("--plan-only", action="store_true", help="skip the encoder taps and decoder evaluations") + ap.add_argument("--evals", nargs="+", type=int, default=[0, 5, 10]) + ap.add_argument("--json", default=None) + args = ap.parse_args() + + import torch + + torch.set_num_threads(4) + from tt_diffusion_planner.device import close_device, open_device + from tt_diffusion_planner.reference.weights import find_weights_dir, load_weights + from tt_diffusion_planner.tt.model import TtDiffusionPlanner + + weights = load_weights(find_weights_dir()) + dev = open_device(allow_fallback=False) + g = dev.compute_with_storage_grid_size() + report = {"grid": f"{g.x}x{g.y}", "precision": args.precision, "scenes": {}} + t0 = time.perf_counter() + from tt_diffusion_planner.tt.config import globs + + opts = {k: globs(v) for k, v in (("ln_fp32", args.ln_fp32), ("hidden_fp32", args.hidden_fp32), + ("split", args.split), ("attn_fp32_acc", args.attn_fp32_acc), + ("attn_matmul", args.attn_matmul)) if v is not None} + tt = TtDiffusionPlanner(dev, weights, debug=True, precision=args.precision, **opts) + report["options"] = tt.build.options() + report["build_s"] = round(time.perf_counter() - t0, 2) + print(f"build {report['build_s']} s, uploaded {tt.build.uploaded_bytes / 2**20:.1f} MiB", flush=True) + runner = tt.runner + from tt_diffusion_planner.tt import inputs as I + + eager_cache = {} + try: + for scene in args.scenes: + with np.load(GOLDENS / f"{scene}.npz", allow_pickle=False) as z: + gold = {k: z[k] for k in z.files if k != "__meta__"} + raw = {k: gold[f"in.{k}"] for k in C.INPUT_NAMES} + prep = hp.prepare(raw, weights.normalization.observation) + res = {} + if args.plan_only: + args.evals = [] + # encoder taps (eager x2) + t1 = time.perf_counter() + taps = tt.encoder_taps(prep, eager=True) + taps2 = taps if args.plan_only else tt.encoder_taps(prep, eager=True) + res["encoder_eager_s"] = round(time.perf_counter() - t1, 2) + res["encoder_deterministic"] = all(np.array_equal(taps[k], taps2[k]) for k in taps) + for name, _ in C.TOKEN_LAYOUT: + rows = np.flatnonzero(gold[f"host.valid.{name}"]) + if rows.size: + res[f"enc.{name}"] = stats(taps[f"enc.{name}"][rows], gold[f"enc.{name}"][rows]) + for c in ("ego", "neighbor", "lane", "route", "polygon", "line_string"): + rows = gold[f"enc.{c}.pre.rows"] + if rows.size: + for part in ("pre", "mixer"): + res[f"enc.{c}.{part}"] = stats(taps[f"enc.{c}.{part}"][rows], gold[f"enc.{c}.{part}"]) + tok = np.flatnonzero(gold["host.token_valid"]) + for name in ["enc.tokens"] + [f"enc.fusion.{i}" for i in range(6)] + ["enc.encoding"]: + res[name] = stats(taps[name][tok], gold[name][tok]) + # decode_once, teacher forced + arows = gold["dec.rows"] + for k in args.evals: + outs = [tt_decode(tt, runner, prep, gold, k) for _ in range(2)] + res[f"dec.eval{k}.deterministic"] = bool(np.array_equal(outs[0], outs[1])) + res[f"dec.eval{k}"] = stats(outs[0][arows][:, 1:], gold["dec.out"][k][:, 1:]) + # plan (eager x2) + t1 = time.perf_counter() + p1 = runner.run_eager("plan", inputs=I.plan_inputs(prep)) + p2 = runner.run_eager("plan", inputs=I.plan_inputs(prep)) + res["plan_eager_s"] = round(time.perf_counter() - t1, 2) + res["plan_deterministic"] = all(np.array_equal(p1[k], p2[k]) for k in p1) + eager_cache[scene] = (prep, p1) + final = p1["final_x0"].reshape(352, 324)[:321].reshape(321, 81, 4) + res["plan.final_x0"] = stats(final[arows], gold["final_x0"][arows]) + res["plan.logit"] = {"dev": [round(float(v), 4) for v in p1["logit"].reshape(-1)[:5]], + "ref": [round(float(v), 4) for v in gold["turn.logit"]]} + out = hp.make_output(final, p1["logit"].reshape(-1)[:5], prep, weights.normalization, + {k: v[3] for k, v in hp.RUNTIME_PARAMS.items()}) + dpos = np.hypot(*(out.poses[:, :2] - gold["out.trajectory"][:, :2]).T) + res["plan.ego_max_m"], res["plan.ego_mean_m"] = round(float(dpos.max()), 4), round(float(dpos.mean()), 4) + res["plan.turn_equal"] = int(out.turn_indicator["command"]) == int(gold["out.turn_command"]) + if gold["out.predicted_agents"].shape[0]: + dxy = out.predicted_agents[..., :2] - gold["out.predicted_agents"][..., :2] + per = np.hypot(*dxy.transpose(2, 0, 1)) + res["plan.nb_median_max_m"] = round(float(np.median(per.max(axis=1))), 4) + report["scenes"][scene] = res + print(json.dumps({scene: res}, indent=1), flush=True) + if not args.no_capture: + t1 = time.perf_counter() + runner.capture() + report["capture_s"] = round(time.perf_counter() - t1, 2) + report["timings_ms"] = runner.timings_ms + for scene, (prep, p1) in eager_cache.items(): + rep = runner("plan", inputs=I.plan_inputs(prep)) + report["scenes"][scene]["plan_replay_equals_eager"] = all(np.array_equal(rep[k], p1[k]) for k in p1) + import ttnn + + ttnn.synchronize_device(dev) + t2 = time.perf_counter() + runner.replay("plan", 5) + ttnn.synchronize_device(dev) + report["scenes"][scene]["plan_replay_ms"] = round((time.perf_counter() - t2) / 5 * 1e3, 3) + print(json.dumps({k: v for k, v in report.items() if k != "scenes"}, indent=1, default=str), flush=True) + print(json.dumps({s: {k: v for k, v in r.items() if k.startswith("plan_")} + for s, r in report["scenes"].items()}, indent=1), flush=True) + finally: + tt.release() + close_device(dev) + if args.json: + Path(args.json).parent.mkdir(parents=True, exist_ok=True) + Path(args.json).write_text(json.dumps(report, indent=1, default=str) + "\n") + return 0 + + +def tt_decode(tt, runner, prep, gold, k): + return tt.decode_once(prep, gold["dec.x_in"][k], float(gold["dec.t"][k]), encoding=gold["enc.encoding"], + eager=True) + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/code/scripts/container_smoke.sh b/code/scripts/container_smoke.sh new file mode 100644 index 0000000000000000000000000000000000000000..4ba7549e68b47df9e7ed3a7046510789ea0fd5ce --- /dev/null +++ b/code/scripts/container_smoke.sh @@ -0,0 +1,137 @@ +#!/usr/bin/env bash +# SPDX-License-Identifier: Apache-2.0 +# Serve ONE serve profile of the BUILT container package, run the smoke test against it, keep the evidence, and always +# stop it -- in one device-lock window. Run it once per serve profile (`tt-model profiles +# /tt_kernel_manifest.json` lists them; without --profile the package's default profile is served): +# +# ROOT=/home/ubuntu/experiments/tt-models +# $ROOT/bin/devrun -t 3600 -k 150 -- env -u HF_TOKEN -u HUGGING_FACE_HUB_TOKEN sg docker -c \ +# "bash code/scripts/container_smoke.sh $ROOT/build/diffusion-planner-p150 [port] [--profile NAME] [--log-dir DIR]" +# +# The smoke test FAILS unless /info reports ETH dispatch and the 12x10 grid and runs what the package pins for the +# profile (dispatch, CQs, variant, weights revision); it also compares the output with the stored CPU reference of the +# sample when one exists (code/tt_diffusion_planner/server/smoke_test.py). Exit code: the smoke test's (0 = PASS); 1 when +# serve fails, 2 on a usage error. +# +# Evidence, kept whatever the outcome (--log-dir, default: logs/smoke/ of this repo, next to code/), named +# [-]-.*: +# .container.log the container's whole log (boot, requests, shutdown), followed while the container stops +# .info.json GET /info as soon as the server is READY (device: dispatch, grid, cores; pins; versions) +# .smoke.json the smoke test's /predict output (the SMOKE_OUT environment variable overrides this path) +# .result.json profile, port, exit code, start / end times +# +# `tt-model serve` returns once the server is READY and leaves the container running, so the stop must happen before +# the lock is released, also when devrun's timeout TERMs this script: the EXIT trap saves the log, then stops the +# container cleanly with SIGTERM (120 s grace, hence devrun -k 150), never `docker kill`, which would leave the chip +# dirty. Needs docker access (sg docker) and python3. +set -u +usage() { echo "usage: container_smoke.sh [port] [--profile NAME] [--log-dir DIR]" >&2; } +PROFILE="" +LOG_DIR="" +POSITIONAL=() +while [ $# -gt 0 ]; do + case "$1" in + --profile) [ $# -ge 2 ] || { usage; exit 2; }; PROFILE="$2"; shift 2 ;; + --profile=*) PROFILE="${1#--profile=}"; shift ;; + --log-dir) [ $# -ge 2 ] || { usage; exit 2; }; LOG_DIR="$2"; shift 2 ;; + --log-dir=*) LOG_DIR="${1#--log-dir=}"; shift ;; + -h|--help) usage; exit 0 ;; + -*) echo "container_smoke.sh: unknown option $1" >&2; usage; exit 2 ;; + *) POSITIONAL+=("$1"); shift ;; + esac +done +[ ${#POSITIONAL[@]} -ge 1 ] && [ ${#POSITIONAL[@]} -le 2 ] || { usage; exit 2; } +STAGED="${POSITIONAL[0]}" +PORT="${POSITIONAL[1]:-20000}" +MANIFEST="$STAGED/tt_kernel_manifest.json" +[ -f "$MANIFEST" ] || { echo "container_smoke.sh: $MANIFEST not found (run tt-model package first)" >&2; exit 2; } +HERE="$(cd "$(dirname "$0")" && pwd)" +PROFILE_ARGS=() +[ -n "$PROFILE" ] && PROFILE_ARGS=(--profile "$PROFILE") + +# The package name and the served profile (tt-model's rule: --profile, else default_profile, else the first serve +# profile), hence the container name tt-model gives it: tt-model-- (tt_kernel/container.py). +read -r NAME PROFILE_NAME < <(python3 - "$MANIFEST" "$PROFILE" <<'EOF' +import json, sys +m = json.load(open(sys.argv[1])) +c = m.get("container") or {} +profiles = c.get("serve_profiles") or [{}] +print(m.get("name") or "model", sys.argv[2] or c.get("default_profile") or profiles[0].get("name") or "default") +EOF +) +[ -n "${NAME:-}" ] || { echo "container_smoke.sh: cannot read the package name from $MANIFEST" >&2; exit 2; } +CONTAINER="tt-model-$NAME-$PROFILE_NAME" + +[ -n "$LOG_DIR" ] || LOG_DIR="$(cd "$HERE/../.." && pwd)/logs/smoke" +if ! mkdir -p "$LOG_DIR" 2>/dev/null || [ ! -w "$LOG_DIR" ]; then + echo "container_smoke.sh: cannot write $LOG_DIR; keeping the evidence in ${TMPDIR:-/tmp}" >&2 + LOG_DIR="${TMPDIR:-/tmp}" +fi +LOG_DIR="$(cd "$LOG_DIR" && pwd)" +STARTED="$(date -u +%Y-%m-%dT%H:%M:%SZ)" +STEM="$LOG_DIR/$NAME${PROFILE:+-$PROFILE}-$(date -u +%Y%m%dT%H%M%SZ)" +SMOKE_JSON="${SMOKE_OUT:-$STEM.smoke.json}" + +fetch() { # fetch URL FILE: one GET with a 30 s timeout; the body goes to FILE + python3 - "$1" "$2" <<'EOF' +import sys, urllib.request +try: + with urllib.request.urlopen(sys.argv[1], timeout=30) as r: + body = r.read() +except Exception as e: # noqa: BLE001 -- best effort: the smoke test reports the server's state + sys.exit(f"container_smoke.sh: GET {sys.argv[1]} failed: {e}") +open(sys.argv[2], "wb").write(body) +EOF +} + +CHILD="" +STOPPED=0 +cleanup() { + local rc=$? + [ "$STOPPED" = 1 ] && return + STOPPED=1 + if [ -n "$CHILD" ]; then kill -TERM "$CHILD" 2>/dev/null; wait "$CHILD" 2>/dev/null; fi + # The log is lost with the container: follow it (whole history, then the shutdown lines) while it stops. + local logger="" + if docker inspect "$CONTAINER" >/dev/null 2>&1; then + docker logs --follow "$CONTAINER" > "$STEM.container.log" 2>&1 & + logger=$! + else + tt-model logs "${PROFILE_ARGS[@]}" "$MANIFEST" > "$STEM.container.log" 2>&1 || true + fi + tt-model stop "${PROFILE_ARGS[@]}" "$MANIFEST" || true + if [ -n "$logger" ]; then + for _ in $(seq 1 30); do kill -0 "$logger" 2>/dev/null || break; sleep 1; done + kill "$logger" 2>/dev/null + wait "$logger" 2>/dev/null + fi + python3 - "$STEM.result.json" "$NAME" "$PROFILE_NAME" "$PORT" "$rc" "$STARTED" "$CONTAINER" "$MANIFEST" <<'EOF' +import datetime, json, sys +out, name, profile, port, rc, started, container, manifest = sys.argv[1:] +ended = datetime.datetime.now(datetime.timezone.utc).strftime("%Y-%m-%dT%H:%M:%SZ") +json.dump({"bundle": name, "profile": profile, "port": int(port) if port.isdigit() else port, "rc": int(rc), + "result": "PASS" if rc == "0" else "FAIL", + "started": started, "ended": ended, "container": container, "manifest": manifest}, open(out, "w"), indent=1) +EOF + echo "container_smoke.sh: rc=$rc; evidence in $STEM.*" +} +trap cleanup EXIT +trap 'exit 143' TERM +trap 'exit 130' INT + +# Each step runs in the background and is waited for: bash defers a trap while a FOREGROUND command runs, so a TERM +# sent to this script alone (devrun's timeout signals the whole process group) would otherwise wait for the step. +step() { + "$@" & + CHILD=$! + wait "$CHILD" + local rc=$? + CHILD="" + return "$rc" +} + +# --port and --profile BEFORE the target: options after it are passed through to the container (tt-model cli rule) +step tt-model serve --port "$PORT" "${PROFILE_ARGS[@]}" "$MANIFEST" || exit 1 +step fetch "http://127.0.0.1:$PORT/info" "$STEM.info.json" +step python3 "$HERE/../tt_diffusion_planner/server/smoke_test.py" --url "http://127.0.0.1:$PORT" --wait 600 \ + --manifest "$MANIFEST" "${PROFILE_ARGS[@]}" --out "$SMOKE_JSON" diff --git a/code/scripts/fetch_samples.sh b/code/scripts/fetch_samples.sh new file mode 100644 index 0000000000000000000000000000000000000000..f4cc0d55494bbfdfb37017d7dbe686a0b6474af4 --- /dev/null +++ b/code/scripts/fetch_samples.sh @@ -0,0 +1,24 @@ +#!/usr/bin/env bash +# SPDX-License-Identifier: Apache-2.0 +# Download the public sample data that diffusion-planner-p150 may NOT redistribute (license unstated or non-commercial) +# into ${DIFFUSION_PLANNER_SAMPLES:-$HOME/.cache/tt_diffusion_planner/samples}, sha256-checked. Tests and benchmarks read it from there. +# Nothing downloaded here is executed; parse it with numpy / sqlite3 only. +set -euo pipefail +DEST="${DIFFUSION_PLANNER_SAMPLES:-$HOME/.cache/tt_diffusion_planner/samples}" +mkdir -p "$DEST" + +fetch() { # fetch + local url="$1" sum="$2" name="$3" + if [ -f "$DEST/$name" ] && echo "$sum $DEST/$name" | sha256sum -c --status; then + echo "ok $name"; return + fi + curl -fL --retry 3 -o "$DEST/$name.part" "$url" + echo "$sum $DEST/$name.part" | sha256sum -c --status || { echo "sha256 mismatch: $name" >&2; exit 1; } + mv "$DEST/$name.part" "$DEST/$name"; echo "fetched $name" +} + +# Example (Autoware demo rosbag; pinned in autoware/ansible/roles/demo_artifacts/tasks/main.yaml:48-53): +# fetch https://autoware-files.s3.us-west-2.amazonaws.com/recordings/bags/demos/sample-rosbag.zip \ +# 5f9d36353393b3d249212153c19049822b1298db56512aa045b4f7f6fc37cf88 sample-rosbag.zip +# nothing to fetch: both shipped samples are redistributable (Apache-2.0) +echo "samples in $DEST" diff --git a/code/scripts/precision_exp.py b/code/scripts/precision_exp.py new file mode 100644 index 0000000000000000000000000000000000000000..dcb0c4edfb0703dcccc3b7364aa689d7a3eec760 --- /dev/null +++ b/code/scripts/precision_exp.py @@ -0,0 +1,147 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Precision experiment for the MLP-Mixer blocks on the device (development tool; run under ``bin/devrun``). + + bin/devrun -t 1200 -- python code/scripts/precision_exp.py [--json logs/diffusion-planner/precision_exp.json] + +Part A: ``ttnn.layer_norm`` vs :func:`tt.layers.layer_norm_fp32` on the real offset-dominated mixer rows (the +reference ``enc.neighbor.pre`` of the samples, block 0 ``norm1``) against float64. Part B: the 6 mixer blocks of ego / +neighbour / lane run alone on the reference's own block input (``enc..pre``, teacher forcing) under several +option sets, compared with the reference ``enc..mixer`` on the valid entities. Eager (no trace). +""" +from __future__ import annotations + +import argparse +import json +import sys +from pathlib import Path + +import numpy as np + +HERE = Path(__file__).resolve() +sys.path.insert(0, str(HERE.parents[1])) + +from tt_diffusion_planner.reference import config as C # noqa: E402 +from tt_diffusion_planner.ttaw.metrics import error_stats # noqa: E402 + +GOLDENS = HERE.parents[4] / "research" / "diffusion-planner" / "goldens" +CONFIGS = { + "default": dict(), + "ln_fp32": dict(ln_fp32=("enc.mixer.*",)), + "hidden_fp32": dict(hidden_fp32=("enc.mixer.*",)), + "ln_hidden_fp32": dict(ln_fp32=("enc.mixer.*",), hidden_fp32=("enc.mixer.*",)), + "all_fp32_w": dict(ln_fp32=("enc.mixer.*",), hidden_fp32=("enc.mixer.*",), + precision="enc.mixer.*=HiFi4+fp32:w=fp32:a=fp32"), + "stream_bf16": dict(precision="enc.mixer.*=HiFi4+fp32:w=bf16:a=bf16"), +} + + +def st(dev, ref): + s = error_stats(np.asarray(dev, np.float64), np.asarray(ref, np.float64)) + return {"pcc": round(s["pcc"], 7), "rel_l2": float(f"{s['rel_l2']:.4g}"), "max_abs": float(f"{s['max_abs']:.4g}")} + + +def main() -> int: + ap = argparse.ArgumentParser() + ap.add_argument("--scenes", nargs="+", default=["kashiwanoha_dense", "straight_road"]) + ap.add_argument("--cats", nargs="+", default=["ego", "neighbor", "lane"]) + ap.add_argument("--configs", nargs="*", default=list(CONFIGS)) + ap.add_argument("--json", default=None) + args = ap.parse_args() + import torch + import ttnn + + torch.set_num_threads(4) + from tt_diffusion_planner.device import close_device, open_device + from tt_diffusion_planner.reference.weights import find_weights_dir, load_weights + from tt_diffusion_planner.tt.encoder import ENTITIES, MixerTrunk + from tt_diffusion_planner.tt.layers import Build, layer_norm_fp32, policy + from tt_diffusion_planner.ttaw.precision import compute_kernel_config + from tt_diffusion_planner.ttaw.tensors import to_device, to_numpy + + p = load_weights(find_weights_dir()).params + gold = {} + for s in args.scenes: + with np.load(GOLDENS / f"{s}.npz", allow_pickle=False) as z: + gold[s] = {k: z[k] for k in z.files if k.startswith(("enc.", "host.valid"))} + dev = open_device(allow_fallback=False) + report = {"ln": {}, "mixer": {}} + try: + # ---- A: LayerNorm accuracy on real rows + g = p["encoder.neighbor_encoder.blocks.0.norm1.gamma"] + b = p["encoder.neighbor_encoder.blocks.0.norm1.beta"] + gd, bd = (to_device(v.reshape(1, 1, 1, -1), dev, "float32") for v in (g, b)) + for s in args.scenes: + x = gold[s]["enc.neighbor.pre"].reshape(1, 1, -1, C.MIXER_CHANNELS).astype(np.float32) + xx = x.astype(np.float64) + ref = (xx - xx.mean(-1, keepdims=True)) / np.sqrt(xx.var(-1, keepdims=True) + 1e-5) * g + b + tx = to_device(x, dev, "float32") + fused = to_numpy(ttnn.layer_norm(tx, epsilon=1e-5, weight=gd, bias=bd, + compute_kernel_config=compute_kernel_config("HiFi4", fp32_acc=True))) + dec = to_numpy(layer_norm_fp32(tx, gd, bd)) + report["ln"][s] = {"fused_fp32": st(fused, ref), "decomposed_fp32": st(dec, ref), + "row_offset_ratio": float(np.abs(xx.mean(-1)).mean() / xx.std(-1).mean())} + print("LN", s, json.dumps(report["ln"][s]), flush=True) + # ---- B: mixer blocks alone, teacher-forced + for name in args.configs: + cfg = dict(CONFIGS[name]) + build = Build(dev, policy(spec=cfg.pop("precision", None)), **cfg) + for cat in args.cats: + trunk = MixerTrunk(build, p, cat) + for s in args.scenes: + rows = gold[s][f"enc.{cat}.pre.rows"] + if rows.size == 0: + continue + x0 = np.zeros((1, ENTITIES[cat], C.MIXER_TOKENS, C.MIXER_CHANNELS), np.float32) + x0[0, rows] = gold[s][f"enc.{cat}.pre"] + out = to_numpy(trunk.mix(to_device(x0, dev, trunk.stream)))[0, rows] + r = st(out, gold[s][f"enc.{cat}.mixer"]) + report["mixer"].setdefault(name, {}).setdefault(cat, {})[s] = r + print(f"MIX {name:15s} {cat:9s} {s:18s} {json.dumps(r)}", flush=True) + del trunk + # ---- C: the island on the device (TF32-like vs split matmuls) -> pre error, and the float64 CPU mixer on + # the device pre (separates the input error from the device mixer) + from tt_diffusion_planner.host import pipeline as hp + from tt_diffusion_planner.reference.model import Encoder, torch_params + from tt_diffusion_planner.tt import inputs as I + + ref_enc = Encoder(torch_params(p, torch.float64)) + weights = load_weights(find_weights_dir()) + + def cpu_mix(x, cat): + N = f"encoder.{cat}_encoder" + x = torch.from_numpy(np.asarray(x, np.float64)) + for i in range(C.MIXER_DEPTH): + B = f"{N}.blocks.{i}" + x = x + ref_enc.mlp(ref_enc.ln(x, f"{B}.norm1").transpose(1, 2), f"{B}.tokens_mlp").transpose(1, 2) + x = x + ref_enc.mlp(ref_enc.ln(x, f"{B}.norm2"), f"{B}.channels_mlp") + return x.numpy() + + preps = {} + for s in args.scenes: + with np.load(GOLDENS / f"{s}.npz", allow_pickle=False) as z: + raw = {k: z[f"in.{k}"] for k in C.INPUT_NAMES} + preps[s] = I.plan_inputs(hp.prepare(raw, weights.normalization.observation)) + for mode, split in (("tf32", ()), ("split", ("enc.island.*",))): + build = Build(dev, policy(), ln_fp32=("enc.mixer.*",), split=split) + for cat in ("ego", "neighbor"): + trunk = MixerTrunk(build, p, cat) + for s in args.scenes: + rows = gold[s][f"enc.{cat}.pre.rows"] + pre_t = trunk.pre(to_device(preps[s][f"{cat}_x"], dev, "float32")) + pre = to_numpy(pre_t)[0, rows] + mix = to_numpy(trunk.mix(pre_t))[0, rows] + r = {"pre": st(pre, gold[s][f"enc.{cat}.pre"]), + "mix_device": st(mix, gold[s][f"enc.{cat}.mixer"]), + "mix_cpu64_on_device_pre": st(cpu_mix(pre, cat), gold[s][f"enc.{cat}.mixer"])} + report.setdefault("island", {}).setdefault(mode, {}).setdefault(cat, {})[s] = r + print(f"ISL {mode:6s} {cat:9s} {s:18s} {json.dumps(r)}", flush=True) + del trunk + finally: + close_device(dev) + if args.json: + Path(args.json).write_text(json.dumps(report, indent=1) + "\n") + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/code/scripts/profile_ops.py b/code/scripts/profile_ops.py new file mode 100644 index 0000000000000000000000000000000000000000..aa248434b732cf5638b60a5fa7296465674eaa98 --- /dev/null +++ b/code/scripts/profile_ops.py @@ -0,0 +1,250 @@ +#!/usr/bin/env python3 +# SPDX-License-Identifier: Apache-2.0 +"""Device profile of diffusion-planner-p150: one eager plan and one traced replay between Tracy signposts. + + ROOT=/home/ubuntu/experiments/tt-models + $ROOT/bin/devrun -t 3600 -- python -m tracy -r -p -v --op-support-count 16000 --no-web-server \\ + -o $ROOT/generated/profiler/diffusion-planner_baseline code/scripts/profile_ops.py + tt-perf-report --start-signpost trace --end-signpost trace_end + python -m tt_diffusion_planner.ttaw.profiling --start trace --end trace_end + +Sections (signposts): ``eager`` .. ``eager_end``, one eager run of the ``plan`` variant after the warm-up +(``TraceRunner.run_eager``: the same graph and programs as the trace, nothing compiled), and ``trace`` .. +``trace_end``, ``--replays`` replays of the captured ``plan`` trace (the served device path) on the uploaded sample. +The device profiler buffer is flushed (``ttnn.ReadDeviceProfiler``) before and after each section. + +Inside the eager section two levels of signposts attribute every op (and, through the identical op order, every +replayed op): + +- ``m:``: the stage of the plan: ``enc..pre`` / ``.mix`` / ``.head`` (the six mixer trunks, then + pool + entity head), ``enc.static``, ``enc.``, ``enc.tokens`` (concat, validity, position + embedding), ``mask`` (key-bias expansion), ``enc.fusion.attn`` / ``.mlp``, ``enc.final_ln``, ``dec.cross_kv`` + (hoisted cross K / V), ``dec.e.preproj``, ``dec.e.b.attn`` / ``.mlp1`` / ``.cross`` / ``.mlp2``, + ``dec.e.final`` (evaluation k = 0..10, DiT block i), ``dec.e.solver`` (the DPM-Solver++(2M) update, the + prefix constraint and the ego-row slice of the next iterate), ``turn``, ``pack``; +- ``c:``: the layer kind: ``split`` (split hi / lo matmul, ``SplitLinear``), ``linear`` (one ``ttnn.linear``), + ``ln32`` (fp32 LayerNorm decomposition), ``ln`` (``ttnn.layer_norm``), ``attn_mm`` (fp32 matmul attention), + ``sdpa``, ``heads`` (head split / merge), ``mask``, and ``glue`` for every op outside a layer (residual adds, + adaLN gates, mixer transposes, solver updates, slices, concats). + +``--no-layer-signposts`` turns both levels off. One plan issues several thousand programs, more than the profiler's +default 1000-program buffer, so run it with ``--op-support-count`` above the program count of the largest section +(the warm-up before the first flush included). The output folder must be absolute and outside the bundle +(PLAN.md 5.1). +""" +from __future__ import annotations + +import argparse +import functools +import json +import time +from pathlib import Path +from typing import Any, Callable, Dict, List + +import numpy as np + +from tt_diffusion_planner import DiffusionPlanner +from tt_diffusion_planner.ttaw.profiling import read_device_profiler, signpost, signposted + +SAMPLE = Path(__file__).resolve().parents[1] / "tt_diffusion_planner" / "samples" / "kashiwanoha_dense.npz" + + +class _Proxy: + """A callable stand-in for a layer object that emits a signpost (``label()``), then calls the layer.""" + + def __init__(self, inner: Any, label: Callable[[], str]): + self.inner, self._label = inner, label + + def __call__(self, *args, **kwargs): + signpost(self._label()) + return self.inner(*args, **kwargs) + + def __getattr__(self, name: str): + return getattr(self.inner, name) + + +def install_signposts(tt) -> Callable[[], None]: + """Signposts on a ``tt.model.TtDiffusionPlanner`` (instance attributes, layer classes and the attention + helpers); returns the function that removes them again.""" + from tt_diffusion_planner.tt import layers as L + from tt_diffusion_planner.tt import model as M + from tt_diffusion_planner.ttaw.ops import attention as A + + undo: List[Callable[[], None]] = [] + state = {"k": -1, "ln": {}} + stack: List[str] = [] + + def set_attr(obj, name, value): + had = name in vars(obj) + old = vars(obj).get(name) + setattr(obj, name, value) + undo.append(lambda: setattr(obj, name, old) if had else delattr(obj, name)) + + def set_item(d, key, value): + old = d[key] + d[key] = value + undo.append(lambda: d.__setitem__(key, old)) + + # ---- layer kinds (class level, with a stack so nested layers restore the outer kind) --------------------- + def kind_wrap(owner, name, kind_of): + fn = getattr(owner, name) + + @functools.wraps(fn) + def inner(*args, **kwargs): + kind = kind_of(*args) + stack.append(kind) + signpost(f"c:{kind}") + try: + return fn(*args, **kwargs) + finally: + stack.pop() + signpost(f"c:{stack[-1] if stack else 'glue'}") + + setattr(owner, name, inner) + undo.append(lambda: setattr(owner, name, fn)) + + kind_wrap(L.Linear, "__call__", lambda *a: "linear") + kind_wrap(L.SplitLinear, "__call__", lambda *a: "split") + kind_wrap(L.LayerNorm, "__call__", lambda self, *a: "ln32" if self.mode == "fp32" else "ln") + kind_wrap(A, "attention_matmul", lambda *a: "attn_mm") + kind_wrap(A, "sdpa", lambda *a: "sdpa") + for name in ("split_qkv", "split_q_kv", "split_heads", "merge_heads"): + kind_wrap(A, name, lambda *a: "heads") + kind_wrap(A, "expand_key_bias", lambda *a: "mask") + + # ---- stages ------------------------------------------------------------------------------------------------ + def method_wrap(obj, name, before=None, after=None): + fn = getattr(obj, name) + + @functools.wraps(fn) + def inner(*args, **kwargs): + if before: + signpost(before()) + out = fn(*args, **kwargs) + if after: + signpost(after()) + return out + + set_attr(obj, name, inner) + + enc, dec = tt.encoder, tt.decoder + for cat, trunk in enc.trunks.items(): + method_wrap(trunk, "pre", before=lambda c=cat: f"m:enc.{c}.pre") + method_wrap(trunk, "mix", before=lambda c=cat: f"m:enc.{c}.mix") + method_wrap(trunk, "pool", before=lambda c=cat: f"m:enc.{c}.head") + set_attr(enc, "static1", _Proxy(enc.static1, lambda: "m:enc.static")) + for cat in list(enc.small): + set_item(enc.small, cat, _Proxy(enc.small[cat], lambda c=cat: f"m:enc.{c}")) + set_attr(enc, "pad_tokens", _Proxy(enc.pad_tokens, lambda: "m:enc.tokens")) + for i, blk in enumerate(enc.blocks): + set_item(blk, "kv", _Proxy(blk["kv"], lambda i=i: f"m:enc.fusion{i}.attn")) + set_item(blk, "n2", _Proxy(blk["n2"], lambda i=i: f"m:enc.fusion{i}.mlp")) + set_attr(enc, "final_norm", _Proxy(enc.final_norm, lambda: "m:enc.final_ln")) + method_wrap(dec, "cross_kv", before=lambda: "m:dec.cross_kv") + method_wrap(dec, "solve", before=lambda: "m:dec.solve") + + def next_eval(): + state["k"] += 1 + state["ln"] = {} + return f"m:dec.e{state['k']}.preproj" + + method_wrap(dec, "evaluate", before=next_eval, after=lambda: f"m:dec.e{state['k']}.solver") + for i, blk in enumerate(dec.blocks): + def ln_label(i=i): + n = state["ln"][i] = state["ln"].get(i, 0) + 1 + return f"m:dec.e{state['k']}.b{i}.{'attn' if n % 2 else 'mlp1'}" + + set_item(blk, "ln", _Proxy(blk["ln"], ln_label)) + set_item(blk, "n3", _Proxy(blk["n3"], lambda i=i: f"m:dec.e{state['k']}.b{i}.cross")) + set_item(blk, "n4", _Proxy(blk["n4"], lambda i=i: f"m:dec.e{state['k']}.b{i}.mlp2")) + set_attr(dec, "fin_ln", _Proxy(dec.fin_ln, lambda: f"m:dec.e{state['k']}.final")) + set_attr(tt, "turn", _Proxy(tt.turn, lambda: "m:turn")) + pack = M.pack_outputs + + def pack_wrap(*args, **kwargs): + signpost("m:pack") + return pack(*args, **kwargs) + + M.pack_outputs = pack_wrap + undo.append(lambda: setattr(M, "pack_outputs", pack)) + # the mask expansion is attributed to its own stage (encoder fusion mask, decoder self-attention mask) + expand = A.expand_key_bias + + def expand_wrap(*args, **kwargs): + signpost("m:mask") + return expand(*args, **kwargs) + + A.expand_key_bias = expand_wrap + undo.append(lambda: setattr(A, "expand_key_bias", expand)) + + def remove() -> None: + for fn in reversed(undo): + fn() + + return remove + + +def main() -> None: + ap = argparse.ArgumentParser(description=__doc__.split("\n")[0]) + ap.add_argument("--input", default=str(SAMPLE)) + ap.add_argument("--dispatch", default=None, choices=["eth", "worker"]) + ap.add_argument("--num-cqs", type=int, default=None, choices=[1, 2]) + ap.add_argument("--replays", type=int, default=1, help="replays inside the trace section") + ap.add_argument("--no-eager", dest="eager", action="store_false") + ap.add_argument("--no-layer-signposts", dest="layers", action="store_false") + ap.add_argument("--json", default=None, help="write the run description (device, trace, timings) here") + a = ap.parse_args() + import ttnn + + from tt_diffusion_planner.host import pipeline as hp + from tt_diffusion_planner.reference import config as C + from tt_diffusion_planner.reference.weights import find_weights_dir + from tt_diffusion_planner.tt import inputs as I + from tt_diffusion_planner.ttaw.io import load_named_arrays + + wd = find_weights_dir() + t0 = time.perf_counter() + with DiffusionPlanner.from_pretrained(dispatch=a.dispatch, num_command_queues=a.num_cqs, + weights_dir=str(wd) if wd else None) as model: + tt, runner, dev = model.tt, model.runner, model.device + print("loaded in %.1f s:" % (time.perf_counter() - t0), json.dumps(model.device_info), flush=True) + raw = load_named_arrays(a.input, C.INPUT_SCHEMA) + inputs = I.plan_inputs(hp.prepare(raw, model.normalization.observation)) + served = model(inputs=raw) # one served plan: upload + replay + readback + ttnn.synchronize_device(dev) + read_device_profiler(dev) # warm-up / capture / first plan out of the buffer + timings: Dict[str, float] = {} + if a.eager: + remove = install_signposts(tt) if a.layers else (lambda: None) + try: + t1 = time.perf_counter() + with signposted("eager"): + eager = runner.run_eager("plan", inputs=inputs) + ttnn.synchronize_device(dev) + timings["eager_ms"] = (time.perf_counter() - t1) * 1e3 + finally: + remove() + read_device_profiler(dev) + final = np.asarray(eager["final_x0"], np.float32) + print("eager final_x0 finite:", bool(np.isfinite(final).all()), flush=True) + runner.upload(inputs) + ttnn.synchronize_device(dev) + read_device_profiler(dev) + t1 = time.perf_counter() + with signposted("trace"): + runner.replay("plan", n=a.replays) + ttnn.synchronize_device(dev) + timings["trace_ms"] = (time.perf_counter() - t1) * 1e3 / a.replays + read_device_profiler(dev) + out = runner.read("plan") + same = bool(np.array_equal(out["final_x0"], eager["final_x0"])) if a.eager else None + desc = {"device": model.device_info, "timings_ms": timings, "replays": a.replays, + "replay_equals_eager": same, "turn_command": int(served.turn_indicator["command"]), + "options": tt.build.options(), "trace": runner.describe()} + print("profile run:", json.dumps(desc, default=str), flush=True) + if a.json: + Path(a.json).write_text(json.dumps(desc, indent=1, default=str) + "\n") + + +if __name__ == "__main__": + main() diff --git a/code/scripts/ref_golden.py b/code/scripts/ref_golden.py new file mode 100644 index 0000000000000000000000000000000000000000..a2a406471e5aeaac53726dc7455bbd33c90ded8d --- /dev/null +++ b/code/scripts/ref_golden.py @@ -0,0 +1,169 @@ +#!/usr/bin/env python3 +# SPDX-License-Identifier: Apache-2.0 +"""Goldens of the fp32 CPU reference (no device): per-module taps and final outputs per scene. + + # research venv (torch + onnx; onnxruntime only for --ort), from the bundle root: + PYTHONPATH=code python code/scripts/ref_golden.py --threads 4 [--ort] \ + [--full-dir ../../research/diffusion-planner/goldens] [--scenes kashiwanoha_dense straight_road ...] + +Scenes: the shipped samples (``code/tt_diffusion_planner/samples/*.npz``: small goldens -> ``tests/goldens``, and the +stored ``/predict`` body of the reference -> ``samples/.reference.json``) and, when the workspace research +directory is present, the seven ORT golden scenes of ``research/diffusion-planner/ort/golden_.npz`` and any +public-dataset scene in ``research/diffusion-planner/public_data/*.npz`` (full goldens only). ``--ort`` also runs ONNX +Runtime on every scene and writes ``/.ort_agreement.json`` (reference vs ORT on the deployed ONNX). +""" +from __future__ import annotations + +import argparse +import hashlib +import json +import sys +import time +from pathlib import Path + +import numpy as np + +CODE = Path(__file__).resolve().parents[1] +if str(CODE) not in sys.path: + sys.path.insert(0, str(CODE)) + +from tt_diffusion_planner.reference import config as C # noqa: E402 +from tt_diffusion_planner.reference.goldens import write_scene # noqa: E402 +from tt_diffusion_planner.reference.pipeline import ReferencePlanner # noqa: E402 +from tt_diffusion_planner.ttaw.metrics import pcc # noqa: E402 + +PKG = CODE / "tt_diffusion_planner" +SAMPLES = PKG / "samples" +SMALL_DIR = PKG / "tests" / "goldens" +RESEARCH = CODE.parents[2] / "research" / "diffusion-planner" + + +def public_ids(dataset: str, which: str): + """Instant ids of ``public_data/inputs/``: ``all``, or ``core`` (the ids the public-data goldens keep + solver iterates for: ``public_data/goldens//core_ids.json``).""" + root = RESEARCH / "public_data" + ids = sorted(p.stem for p in (root / "inputs" / dataset).glob("*.npz")) + if which == "core": + core = root / "goldens" / dataset / "core_ids.json" + keep = set(json.loads(core.read_text()).get("core_ids_with_denoising_steps", [])) if core.is_file() else set() + ids = [i for i in ids if i in keep] + return ids + + +def scene_sources(selected, public=None, public_which="core"): + """``{scene: (raw source, kind)}``: shipped samples first, then research ORT scenes, then (``--public``) the + public-data instants of one dataset (nuScenes-derived: CC BY-NC-SA, local goldens only).""" + out = {} + if public is None: + for p in sorted(SAMPLES.glob("*.npz")): + out[p.stem] = (p, "sample") + for p in sorted((RESEARCH / "ort").glob("golden_*.npz")): + out.setdefault(p.stem[len("golden_"):], (p, "research")) + else: + for i in public_ids(public, public_which): + out[f"public/{public}/{i}"] = (RESEARCH / "public_data" / "inputs" / public / f"{i}.npz", "public") + if selected: + missing = sorted(set(selected) - set(out)) + if missing: + raise SystemExit(f"unknown scenes {missing}; have {sorted(out)}") + out = {k: v for k, v in out.items() if k in selected} + return out + + +def load_raw(path: Path, kind: str): + with np.load(path, allow_pickle=False) as z: + if kind in ("research", "public"): # dp_reference.py / nuscenes_dp.py layout: raw/ + return {k: np.asarray(z["raw/" + k], np.float32) for k in C.INPUT_NAMES} + return {k: np.asarray(z[k], np.float32) for k in C.INPUT_NAMES} + + +def ort_agreement(weights_dir: Path, raw, g, threads: int): + from tt_diffusion_planner.reference.ort import OrtPlanner + + o = OrtPlanner(weights_dir, threads=threads).run(raw) + rows = g["dec.rows"] + enc_rows = np.flatnonzero(g["host.token_valid"]) + return {"encoding_pcc_valid": pcc(g["enc.encoding"][enc_rows], o.encoding[0][enc_rows]), + "encoding_max_abs": float(np.abs(g["enc.encoding"] - o.encoding[0]).max()), + "final_x0_pcc_valid": pcc(g["final_x0"][rows], o.final_x0[0][rows]), + "final_x0_max_abs_valid": float(np.abs(g["final_x0"][rows] - o.final_x0[0][rows]).max()), + "logit_max_abs": float(np.abs(g["turn.logit"] - o.logit[0]).max()), + "turn_logit_ref": g["turn.logit"].tolist(), "turn_logit_ort": o.logit[0].tolist()} + + +def stored_ort_agreement(g, golden: Path): + """Reference vs a stored ORT golden in the ``dp_reference.py`` layout (``encoding``, ``final_x_normalized``, + ``logit_multi``): the public-data goldens of ``research/diffusion-planner/public_data/goldens/``.""" + with np.load(golden, allow_pickle=False) as z: + enc, fx, lg = z["encoding"][0], z["final_x_normalized"][0], z["logit_multi"][0] + rows = g["dec.rows"] + tok = np.flatnonzero(g["host.token_valid"]) + return {"ort_golden": str(golden), + "encoding_pcc_valid": pcc(g["enc.encoding"][tok], enc[tok]), + "encoding_max_abs": float(np.abs(g["enc.encoding"][tok] - enc[tok]).max()), + "final_x0_pcc_valid": pcc(g["final_x0"][rows], fx[rows]), + "final_x0_max_abs_valid": float(np.abs(g["final_x0"][rows] - fx[rows]).max()), + "logit_max_abs": float(np.abs(g["turn.logit"] - lg).max())} + + +def main(argv=None) -> int: + ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) + ap.add_argument("--weights-dir", default=None) + ap.add_argument("--full-dir", type=Path, default=RESEARCH / "goldens", + help="where the large per-scene goldens go (never inside the bundle)") + ap.add_argument("--scenes", nargs="*", default=None) + ap.add_argument("--threads", type=int, default=4) + ap.add_argument("--ort", action="store_true", help="also compare with ONNX Runtime (research venv)") + ap.add_argument("--no-reference-json", action="store_true") + ap.add_argument("--public", default=None, help="a public_data/inputs/ (e.g. nuscenes): its instants " + "instead of the samples / research scenes, as lite goldens") + ap.add_argument("--public-ids", default="core", choices=["core", "all"]) + a = ap.parse_args(argv) + if a.full_dir.resolve().is_relative_to(CODE.parent.resolve()): + raise SystemExit("--full-dir must be outside the bundle: the full goldens are 10-25 MB per scene") + ref = ReferencePlanner(a.weights_dir, threads=a.threads) + report, seen = {}, {} + for scene, (path, kind) in scene_sources(a.scenes, a.public, a.public_ids).items(): + t0 = time.perf_counter() + raw = load_raw(path, kind) + digest = hashlib.sha256(b"".join(raw[k].tobytes() for k in C.INPUT_NAMES)).hexdigest() + if digest in seen: # e.g. research golden_straight == the shipped straight_road sample + report[scene] = {"same_inputs_as": seen[digest]} + print(scene, "skipped: same inputs as", seen[digest], flush=True) + continue + seen[digest] = scene + small = SMALL_DIR if kind == "sample" else None + r = write_scene(ref, raw, scene, a.full_dir, small, meta={"source": str(path), "kind": kind}, + lite=(kind == "public")) + entry = {"paths": r["paths"], "valid_counts": r["info"]["valid_counts"]} + if kind == "sample" and not a.no_reference_json: + body = ref(inputs=raw).to_dict() + body["timing_ms"] = {} + ref_path = SAMPLES / f"{scene}.reference.json" + ref_path.write_text(json.dumps(body, indent=1) + "\n") + entry["reference_json"] = str(ref_path) + stored = None + if kind == "public": + stored = RESEARCH / "public_data" / "goldens" / a.public / f"golden_{path.stem}.npz" + if stored is not None and stored.is_file(): # the dataset work already ran ORT on this instant + agree = stored_ort_agreement(r["goldens"], stored) + (a.full_dir / f"{scene}.ort_agreement.json").write_text(json.dumps(agree, indent=1) + "\n") + entry["ort"] = {k: v for k, v in agree.items() if k != "ort_golden"} + elif a.ort: + agree = ort_agreement(ref.weights.path, raw, r["goldens"], a.threads) + a.full_dir.mkdir(parents=True, exist_ok=True) + (a.full_dir / f"{scene}.ort_agreement.json").write_text(json.dumps(agree, indent=1) + "\n") + entry["ort"] = {k: v for k, v in agree.items() if not k.startswith("turn_logit")} + entry["seconds"] = round(time.perf_counter() - t0, 1) + report[scene] = entry + print(scene, json.dumps(entry), flush=True) + a.full_dir.mkdir(parents=True, exist_ok=True) + index = a.full_dir / "index.json" + merged = json.loads(index.read_text()) if index.is_file() else {} + merged.update(report) + index.write_text(json.dumps(merged, indent=1, sort_keys=True) + "\n") + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/code/scripts/split_error.py b/code/scripts/split_error.py new file mode 100644 index 0000000000000000000000000000000000000000..a8829ea9146cba2d67d7789785eec7e6b6efcfef --- /dev/null +++ b/code/scripts/split_error.py @@ -0,0 +1,124 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Where does the end-to-end error of a scene come from? (development tool; device, run under ``bin/devrun``) + + bin/devrun -t 1200 -- python code/scripts/split_error.py --public scene-0103_kf14 --scenes straight_road + +For each scene the ego / neighbour displacement vs the fp32 CPU reference of: + +- ``device``: the full device plan; +- ``enc_only``: the device encoding (``encoder_taps``) + the CPU fp32 decoder and solver (encoder error alone); +- ``dec_only``: the CPU encoding + the device decoder (``decode_once`` replays) driven by the host solver (decoder + error alone, including its compounding over the 11 evaluations). +""" +from __future__ import annotations + +import argparse +import json +import sys +from pathlib import Path + +import numpy as np + +HERE = Path(__file__).resolve() +sys.path.insert(0, str(HERE.parents[1])) + +from tt_diffusion_planner.host import pipeline as hp # noqa: E402 +from tt_diffusion_planner.host.solver import apply_prefix_constraint, dpm_solver_sample # noqa: E402 +from tt_diffusion_planner.reference import config as C # noqa: E402 + +RES = HERE.parents[4] / "research" / "diffusion-planner" + + +def load_raw(scene: str, public: bool): + if public: + with np.load(RES / "public_data" / "inputs" / "nuscenes" / f"{scene}.npz", allow_pickle=False) as z: + return {k: np.array(z[f"raw/{k}"]) for k in C.INPUT_NAMES} + with np.load(RES / "goldens" / f"{scene}.npz", allow_pickle=False) as z: + return {k: np.array(z[f"in.{k}"]) for k in C.INPUT_NAMES} + + +def main() -> int: + ap = argparse.ArgumentParser() + ap.add_argument("--public", nargs="*", default=[]) + ap.add_argument("--scenes", nargs="*", default=[]) + ap.add_argument("--json", default=None) + for opt in ("ln-fp32", "hidden-fp32", "split", "attn-fp32-acc", "attn-matmul"): + ap.add_argument(f"--{opt}", default=None, help="module globs (comma-separated); default: the knob") + args = ap.parse_args() + from tt_diffusion_planner.tt.config import globs + + opts = {k: globs(getattr(args, k)) for k in ("ln_fp32", "hidden_fp32", "split", "attn_fp32_acc", "attn_matmul") + if getattr(args, k) is not None} + import torch + + torch.set_num_threads(4) + from tt_diffusion_planner.device import close_device, open_device + from tt_diffusion_planner.reference.model import Decoder, Encoder, torch_params + from tt_diffusion_planner.reference.weights import find_weights_dir, load_weights + from tt_diffusion_planner.tt.model import TtDiffusionPlanner + + w = load_weights(find_weights_dir()) + params = {k: v[3] for k, v in hp.RUNTIME_PARAMS.items()} + P = torch_params(w.params) + enc_ref, dec_ref = Encoder(P), Decoder(P) + dev = open_device(allow_fallback=False) + report = {} + try: + tt = TtDiffusionPlanner(dev, w, debug=True, **opts) + report["options"] = tt.build.options() + tt.capture() + for scene, public in [(s, True) for s in args.public] + [(s, False) for s in args.scenes]: + prep = hp.prepare(load_raw(scene, public), w.normalization.observation) + cs = prep.decoder.current_states + + def solve(model_fn): + res = dpm_solver_sample(prep.x_T, model_fn, lambda x: apply_prefix_constraint(x, cs)) + return res.final_x + + def cpu_plan(encoding): + kv = dec_ref.cross_kv(torch.from_numpy(np.asarray(encoding, np.float32))) + with torch.no_grad(): + return solve(lambda x, t: dec_ref.forward(x, t, kv, prep.decoder.agent_valid).numpy()) + + def out(final_x0): + o = hp.make_output(final_x0, np.zeros(5, np.float32), prep, w.normalization, params) + return o.poses[:, :2], o.predicted_agents[..., :2] + + with torch.no_grad(): + enc = enc_ref.forward(prep.features).numpy() + ref_ego, ref_nb = out(cpu_plan(enc)) + dev_enc = tt.encoder_taps(prep)["enc.encoding"] + cases = { + "device": tt.forward(prep)["final_x0"], + "enc_only": cpu_plan(dev_enc), + "dec_only": solve(lambda x, t: _decode(tt, prep, x, t, enc)), + } + res = {} + for name, fx in cases.items(): + ego, nb = out(fx) + d = np.hypot(*(ego - ref_ego).T) + r = {"ego_max_m": round(float(d.max()), 4), "ego_mean_m": round(float(d.mean()), 4)} + if nb.shape[0]: + per_agent = np.hypot(*(nb - ref_nb).transpose(2, 0, 1)).max(1) + r["nb_median_max_m"] = round(float(np.median(per_agent)), 4) + res[name] = r + res["encoding_pcc"] = float(np.corrcoef(dev_enc[prep.features.token_valid].ravel(), + enc[prep.features.token_valid].ravel())[0, 1]) + report[scene] = res + print(scene, json.dumps(res), flush=True) + tt.release() + finally: + close_device(dev) + if args.json: + Path(args.json).write_text(json.dumps(report, indent=1) + "\n") + return 0 + + +def _decode(tt, prep, x, t, enc): + out = tt.decode_once(prep, x, float(t), encoding=enc) + out[:, 0] = 0.0 # the t = 0 slot (masked on the device) is overwritten by the prefix constraint anyway + return out + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/code/tt_diffusion_planner/__init__.py b/code/tt_diffusion_planner/__init__.py new file mode 100644 index 0000000000000000000000000000000000000000..4cac9c79b979699aee3fb1a07e7d6fdb312a97f0 --- /dev/null +++ b/code/tt_diffusion_planner/__init__.py @@ -0,0 +1,41 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Diffusion Planner v5.0 (Autoware diffusion_planner) on a Tenstorrent Blackhole p150 (tt-nn), packaged as +changh95/diffusion-planner-p150. + +Python API (see PYTHON.md):: + + from tt_diffusion_planner import DiffusionPlanner + + with DiffusionPlanner.from_pretrained(device_id=0) as model: + out = model(inputs="code/tt_diffusion_planner/samples/kashiwanoha_dense.npz") + +Importing the package has no side effects (no device, no ttnn / torch import, no network); the names below load +on first use. The HTTP server is ``tt_diffusion_planner.server.app:app`` (SERVING.md). ``tt_diffusion_planner.ttaw`` +is the vendored shared package of the Autoware ports (``ttaw/VENDORED.json`` records its version and file hashes; +never edit it here). ``tt_diffusion_planner.reference`` is the fp32 CPU reference (torch) and +``tt_diffusion_planner.host`` the node's pre- and post-processing (numpy). +""" + +__version__ = "0.1.0" +__all__ = ["DiffusionPlanner", "Output", "open_device", "load_inputs", "INPUT_SCHEMA", "__version__"] + +_LAZY = { + "DiffusionPlanner": (".api", "DiffusionPlanner"), + "Output": (".api", "Output"), + "open_device": (".device", "open_device"), + "load_inputs": (".io", "load_inputs"), + "INPUT_SCHEMA": (".reference.config", "INPUT_SCHEMA"), +} + + +def __getattr__(name): + if name in _LAZY: + import importlib + + module, attr = _LAZY[name] + return getattr(importlib.import_module(module, __name__), attr) + raise AttributeError(f"module {__name__!r} has no attribute {name!r}") + + +def __dir__(): + return sorted(list(globals()) + __all__) diff --git a/code/tt_diffusion_planner/api.py b/code/tt_diffusion_planner/api.py new file mode 100644 index 0000000000000000000000000000000000000000..d26269fe770b063b3927a82d3d611b0bf0fb9137 --- /dev/null +++ b/code/tt_diffusion_planner/api.py @@ -0,0 +1,125 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Python API: Diffusion Planner v5.0 (Autoware diffusion_planner) on one Tenstorrent Blackhole p150. + + from tt_diffusion_planner import DiffusionPlanner + + with DiffusionPlanner.from_pretrained(device_id=0) as model: # weights -> HF cache, device open, traces captured + out = model(inputs="code/tt_diffusion_planner/samples/kashiwanoha_dense.npz") + print(out.to_dict()) # the same JSON as POST /predict + +The contract shared by every bundle of the Autoware collection is the vendored ``ttaw.api_base.ModelBase`` +(BUNDLE_CONVENTIONS.md section 8): ``from_pretrained`` resolves the pinned weights before it claims the chip, opens it +(ETH dispatch, 12x10), builds the graph and captures every trace variant in ``warmup_variants``, so the first call is +as fast as the later ones; calls are serialised by a lock (one chip, batch 1) and fill ``timing_ms``; ``close()`` is +idempotent, also runs at interpreter exit, and closes the chip only if the model opened it. The HTTP server +(``tt_diffusion_planner.server.app``) calls this class, so ``/predict`` and ``model(...)`` agree bit for bit. + +Input: ``inputs=`` holds the 15 raw tensors of the node's ``DiffusionPlannerCore::create_input_data`` (ego frame, +before normalization, batch 1; ``INPUT_SCHEMA``). Converting ROS messages and the Lanelet2 map into them, and the +node's temporal state (agent buffers, ego history, RTC prefix of ``sampled_trajectories``), stay with the client. +The hooks delegate to ``tt_diffusion_planner.host`` (the node's pre- and post-processing, numpy) around +``tt_diffusion_planner.tt`` (the ttnn graph: encoder + 11 DiT evaluations with the DPM-Solver++(2M) update + turn head, +run by a ``ttaw.trace.TraceRunner``). Importing this module has no side effects (ttnn / torch only inside hooks). +""" +from __future__ import annotations + +from typing import Any, Dict + +from . import io as tio +from .device import DEVICE_DEFAULTS +from .host import pipeline as hp +from .reference import config as C +from .ttaw.api_base import ModelBase +from .ttaw.outputs import Trajectory + +__all__ = ["DiffusionPlanner", "Output"] + +# The result class of this model (ttaw.outputs: Detections3D, Detections2D, Segmentation3D, Mask2D, Trajectory). +Output = Trajectory + + +class DiffusionPlanner(ModelBase): + """Diffusion Planner v5.0 (Autoware diffusion_planner) on one Blackhole p150. Create it with + :meth:`from_pretrained`.""" + + MODEL_NAME = "diffusion-planner-p150" + ENV_PREFIX = "DIFFUSION_PLANNER" # prefix of the environment knobs (SERVING.md section 3.4) + DEFAULT_REPO = "AutowareFoundation/diffusion_planner" + DEFAULT_TAG = "v5.0" # the Autoware ansible artifacts pin; the node loads only major version 5 + # the commit DEFAULT_TAG points to (pinned: tags can move) + DEFAULT_REVISION = "423efde67f5414734da43a7ad856c17ceb8b51aa" + ALLOW_PATTERNS = ["diffusion_planner_encoder.onnx", "diffusion_planner_decoder.onnx", + "diffusion_planner_turn_indicator.onnx", "diffusion_planner.param.json"] + VARIANTS = ["default"] # load-time: the multi-step graph with dpm_solver_steps = 10 + DEFAULT_VARIANT = "default" + INPUT_KIND = "planner" # lidar | camera | multicam | lidar+multicam | planner + CAMERA_ORDER = () + POINT_FIELDS = tio.DEFAULT_POINT_FIELDS # () : no point cloud + # turn-indicator logit order (dimensions.hpp:74-79); the published command is the index for 0..3 + LABELS = C.TURN_INDICATOR_LABELS + # Per-request knobs: name -> (type, min, max, default); host-side post-processing only (the node's YAML defaults). + RUNTIME_PARAMS = hp.RUNTIME_PARAMS + EXTRA_INPUTS = () + # the ONNX-named raw tensors (name -> (shape, dtype)), decoded and checked on every call (API and server alike) + INPUT_SCHEMA = C.INPUT_SCHEMA + DEVICE_DEFAULTS = DEVICE_DEFAULTS # validated open parameters (device.py) + + # ---- port-specific hooks (called by ModelBase; keep host work out of _forward) --------------------------- + def _build(self) -> None: + """Weights (the three ONNX files + param JSON, read as data by ``reference.weights``) -> the ttnn graph of + ``tt_diffusion_planner.tt`` registered as the variants of a ``ttaw.trace.TraceRunner`` (persistent inputs, + RT-dev solver / adaLN tables and states allocated here, before any capture). No capture here.""" + from .reference.weights import load_weights + from .tt.model import TtDiffusionPlanner + + unknown = sorted(set(self.compile_params) - {"precision"}) + if unknown: + raise TypeError(f"unknown compile parameter(s) {unknown}; allowed: precision (extra precision rules, " + "e.g. 'dec.*=HiFi2+fp32'); the LN_FP32 / HIDDEN_FP32 options are DIFFUSION_PLANNER_* knobs") + self.planner_weights = load_weights(self.weights_path) + self.normalization = self.planner_weights.normalization + self.tt = TtDiffusionPlanner(self.device, self.planner_weights, precision=self.compile_params.get("precision")) + self.runner = self.tt.runner + + def _warm_one(self, variant: Dict[str, Any]) -> None: + """``TraceRunner.capture`` warms every pending variant eagerly (kernel JIT, program cache) before any capture, + then captures with program-cache misses forbidden; idempotent.""" + self.runner.capture() + + def _prepare(self, points: Any = None, inputs: Any = None, **other: Any) -> hp.Prepared: + """The node's host pre-processing (``host.prepare``): normalization (all-zero rows kept), speed masks, the + encoder's host features, the decoder masks and the solver's initial state.""" + given = sorted(k for k, v in {"points": points, **other}.items() if v is not None) + if given: + raise tio.InputError(f"this model takes only `inputs` (the planner tensors), not {given}") + if inputs is None: + raise tio.InputError("this model needs `inputs`: the 15 planner tensors of INPUT_SCHEMA") + return hp.prepare(inputs, self.normalization.observation) + + def _forward(self, prepared: hp.Prepared) -> Dict[str, Any]: + """Upload the host features into the persistent device inputs, replay the plan's trace(s) and read back + ``final_x0`` (normalised [321, 81, 4]), the turn logits and, when asked, the solver iterates.""" + return self.tt.forward(prepared) + + def _postprocess(self, raw: Dict[str, Any], prepared: hp.Prepared, params: Dict[str, Any]) -> Output: + """The node's post-processing (``host.make_output``): denormalisation, trajectory velocity / force-stop / + acceleration, predicted neighbour paths, turn-indicator decision.""" + return hp.make_output(raw["final_x0"], raw["logit"], prepared, self.normalization, params, + model=self.MODEL_NAME, denoising_steps=raw.get("denoising_steps")) + + def _release(self) -> None: + """Release the traces and persistent device tensors; also called when ``from_pretrained`` fails half-way.""" + runner = getattr(self, "runner", None) + if runner is not None: + runner.release() + + def extra_info(self) -> Dict[str, Any]: + """Additions to ``model.info`` and ``/info``: the trace variants, CQs and persistent tensors of the runner.""" + runner = getattr(self, "runner", None) + info: Dict[str, Any] = {"dpm_solver_steps": C.DPM_SOLVER_STEPS, "input_names": list(C.INPUT_NAMES)} + tt = getattr(self, "tt", None) + if tt is not None: + info.update(tt.describe()) + elif runner is not None: + info["trace"] = runner.describe() + return info diff --git a/code/tt_diffusion_planner/calib/README.md b/code/tt_diffusion_planner/calib/README.md new file mode 100644 index 0000000000000000000000000000000000000000..8fdf0a61d28e1b1d92c8f59d6fd751a1b464aa55 --- /dev/null +++ b/code/tt_diffusion_planner/calib/README.md @@ -0,0 +1,6 @@ +# Calibration presets + +None: the planner takes no calibration. Its inputs are tensors in the ego (`base_link`) frame that the client builds +from ROS messages and the Lanelet2 map (README "Quickstart", SERVING.md 3.1), so there is no camera or LiDAR extrinsic to +send. The directory exists because the shared server of the Autoware ports lists `*.json` presets from it in `/info` +(`calibration_presets`, empty here); a request that carries `calibration` is refused (HTTP 400). diff --git a/code/tt_diffusion_planner/device.py b/code/tt_diffusion_planner/device.py new file mode 100644 index 0000000000000000000000000000000000000000..57250c7e7f07ac75ae1f5b9c2aa1c158c83223fb --- /dev/null +++ b/code/tt_diffusion_planner/device.py @@ -0,0 +1,61 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Opening the Blackhole p150 the way every published number of diffusion-planner-p150 was measured. + +The implementation is the vendored ``ttaw.device`` (C01): dispatch on the idle ETH cores +(``ttnn.DispatchCoreConfig(ttnn.DispatchCoreType.ETH)``), which gives a 12x10 = 120-core compute grid on a p150 +(11x10 with WORKER dispatch, the A/B switch). ETH dispatch needs ``patches/tt-metal-eth-dispatch.patch`` on tt-metal +44d6650 (the container image is built from a patched tree); if the ETH open fails, a ``RuntimeWarning`` is issued and +WORKER dispatch is used, unless ``allow_fallback=False``. Never hard-code the grid: use :func:`compute_grid` or +``device.compute_with_storage_grid_size()``. + +This module binds it to the validated open parameters of this port, :data:`DEVICE_DEFAULTS` (the model class uses +the same dict), which the ``DIFFUSION_PLANNER_DISPATCH``, ``_NUM_CQS``, ``_L1_SMALL``, ``_TRACE_REGION`` and +``_WORKER_L1_SIZE`` variables and ``TT_DEVICE_ID`` override per process (SERVING.md section 3.4). Importing it has +no side effects (``ttnn`` is imported when a device is opened). +""" +from __future__ import annotations + +import contextlib +import dataclasses +from typing import Any, Dict, Iterator, Optional + +from .ttaw.device import DeviceConfig, close_device, compute_grid, core_grid, describe_device, full_core_range_set + +__all__ = ["ENV_PREFIX", "DEVICE_DEFAULTS", "DeviceConfig", "device_config", "open_device", "device_session", + "close_device", "describe_device", "compute_grid", "core_grid", "full_core_range_set"] + +ENV_PREFIX = "DIFFUSION_PLANNER" + +# Validated device-open parameters of this port (fill per model; the card's numbers are measured with them). +DEVICE_DEFAULTS: Dict[str, Any] = { + "num_command_queues": 1, # 1, or 2 when the input upload (CQ1) overlaps the trace (CQ0) + "l1_small_size": 32768, # L1_SMALL bytes per core (conv / pool config tensors) + # DRAM bytes for the traces: measured 74.6 MB for the plan, 96.9 MB with the tests' debug variants (PORT_LOG 5.3) + "trace_region_size": 192 << 20, +} + + +def device_config(**overrides: Any) -> DeviceConfig: + """:data:`DEVICE_DEFAULTS` < the ``DIFFUSION_PLANNER_*`` / ``TT_DEVICE_ID`` environment < explicit non-None + ``overrides`` + (``device_id``, ``dispatch``, ``num_command_queues``, ``l1_small_size``, ``trace_region_size``, + ``worker_l1_size``, ``allow_fallback``): the same resolution as ``DiffusionPlanner.from_pretrained``.""" + config = DeviceConfig.from_env(ENV_PREFIX, **DEVICE_DEFAULTS) + return dataclasses.replace(config, **{k: v for k, v in overrides.items() if v is not None}) + + +def open_device(device_id: Optional[int] = None, *, dispatch: Optional[str] = None, **overrides: Any): + """Open one chip like the published numbers: ETH dispatch (``dispatch="worker"`` is the A/B opt-in) and this + port's sizes. Close it with :func:`close_device`, or use :func:`device_session`.""" + return device_config(device_id=device_id, dispatch=dispatch, **overrides).open() + + +@contextlib.contextmanager +def device_session(device_id: Optional[int] = None, **overrides: Any) -> Iterator[Any]: + """``with device_session() as dev:`` opens with :func:`open_device` and always closes, also when the body + raises.""" + device = open_device(device_id, **overrides) + try: + yield device + finally: + close_device(device) diff --git a/code/tt_diffusion_planner/host/__init__.py b/code/tt_diffusion_planner/host/__init__.py new file mode 100644 index 0000000000000000000000000000000000000000..9b7a4ff4204ce770a5addfa523ff7c84148afdb0 --- /dev/null +++ b/code/tt_diffusion_planner/host/__init__.py @@ -0,0 +1,27 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Host pre- and post-processing of the Autoware Diffusion Planner (numpy only, no ttnn, no torch). + +Exact ports of the node's host code (autoware_universe @ 9ceaccf ``planning/autoware_diffusion_planner``) and of +the in-graph pre-processing of the encoder that the TT port moves to the host (PLAN.md section 2.12 "Host fallbacks": +normalization, position features, masks, post-processing): + +- :mod:`.normalize` input normalization that keeps all-zero rows at zero, speed-limit masks + (``preprocessing_utils.cpp:34-84``, ``inference/utils.hpp:112-123``); +- :mod:`.features` what the encoder graph computes before its first matmul: history truncation, validity masks, + the neighbour velocity zeroing and valid-step flag, lane attributes and speed selection, polygon / line-string + deltas, the 14-dim position features (ONNX ``atan2`` decomposition and the polygon / line-string pseudo-heading + quirk), the fusion key mask; decoder agent mask and current states (SPEC 3.8, 4.3); +- :mod:`.solver` DPM-Solver++(2M) with denoise-to-zero, its float32 scalar schedule computed with the C library + like ``dpm_solver.cpp``, the prefix constraint (``multi_step_inference.cpp:300-340``); +- :mod:`.postprocess` denormalization, poses with Eigen's quaternion of the unnormalised rotation, the trajectory + velocity / force-stop / acceleration rules, predicted neighbour paths, the turn-indicator decision + (``postprocessing_utils.cpp``, ``turn_indicator_manager.cpp``); +- :mod:`.pipeline` ``prepare(raw) -> Prepared`` and ``make_output(...) -> Trajectory``: the two halves the Python + API, the HTTP server and the CPU reference share around the network. + +Importing this package has no side effects. +""" +from .normalize import normalize_inputs, speed_masks # noqa: F401 +from .pipeline import Prepared, make_output, prepare # noqa: F401 + +__all__ = ["normalize_inputs", "speed_masks", "Prepared", "prepare", "make_output"] diff --git a/code/tt_diffusion_planner/host/features.py b/code/tt_diffusion_planner/host/features.py new file mode 100644 index 0000000000000000000000000000000000000000..458b0373a7548ed49ef57fb4df0a270774b806cd --- /dev/null +++ b/code/tt_diffusion_planner/host/features.py @@ -0,0 +1,203 @@ +# SPDX-License-Identifier: Apache-2.0 +"""The encoder's in-graph pre-processing and the decoder's masks, computed on the host from normalised inputs. + +Everything the v5.0 encoder graph does before its first matmul is tensor plumbing (slices, ``!= 0`` tests, OneHot, +``Atan`` with quadrant ``Where`` s): the TT port computes it here once per plan and uploads the results (PLAN.md 2.12; +SPEC 3.8, 8.2.6: padded lanes produce ``0/0`` headings that only a ``Where`` discards, which an arithmetic device +``where`` would propagate). The CPU reference consumes the same arrays, so the reference-vs-ONNX Runtime test also +proves this module against the graph (``/encoder/*/Where*``, ``/encoder/Concat_4`` / ``Concat_5`` taps). + +Semantics (T4M ``model/module/encoder.py``; checked against the exported graph): + +- ego: only the 6 OLDEST history rows are kept (rows 6..30 zeroed); the ego token is always valid and its position + feature is the last row of the truncated history, i.e. zeros; +- neighbours: rows 0..24 zeroed; a step is valid if any of dims 0..7 is non-zero, an agent if any step is; the type + one-hot (dims 8..10) and the position feature come from the last step; velocities (dims 4, 5) are zeroed AFTER the + validity test and a valid-step flag is appended (9 channels); invalid agents are zeroed; +- static objects: valid if any of the 10 values is non-zero (always zeros in Autoware, so never valid); +- lanes / route: dims 0..7 per point (zeroed for invalid lanes), attributes = dims 8..32 of point 0, speed limit and + its mask, position feature = point 10 with heading ``atan2(dy, dx)`` exported as ``Atan(dy / dx)`` plus quadrant + ``Where`` s; +- polygons / line strings: ``[x, y, type one-hot..., dx, dy]`` with dx, dy = next point minus point (0 for the last + point), valid if any of the first FOUR columns is non-zero, position feature = point 20 / 10 with the NON-geometric + heading ``atan2(col3, col2)``: atan2(dx, is_intersection_area) for polygons, atan2(is_road_border, is_stop_line) + for line strings (SPEC 3.8: port the quirk exactly); +- goal / ego shape / turn indicators: always valid; positions (goal) and (0, 0, 1, 0); turn indicators drop the + current report (``[:, :-1]``, 30 values); the turn token reuses the ego-shape class id 8; +- fusion key mask: invalid tokens are masked with -inf, the ego key is forced valid (``encoder.py:833``); +- decoder: a neighbour is a valid attention key if its current state (``neighbor_agents_past[:, 30, :4]``) is not + all zero; current states = (ego_current_state[:4], neighbour current states) for the prefix constraint. +""" +from __future__ import annotations + +from dataclasses import dataclass +from typing import Dict, Mapping + +import numpy as np + +from ..reference import config as C + +__all__ = ["EncoderFeatures", "DecoderMasks", "encoder_features", "decoder_masks", "atan2_onnx", "line_features"] + +PI_F32 = np.float32(3.1415927) # the ONNX constant of the exported atan2 (``/encoder/lane_encoder/Constant_13``) + + +def atan2_onnx(y: np.ndarray, x: np.ndarray) -> np.ndarray: + """``torch.atan2`` as exported to ONNX opset 20: ``a = Atan(y / x)``; ``Where(x < 0, Where(y > 0, a + pi, a - pi), + a)`` (float32). Differs from IEEE atan2 only on signed zeros (x = -0, y = 0 -> -pi) and 0/0 (NaN, masked later).""" + y = np.asarray(y, np.float32) + x = np.asarray(x, np.float32) + with np.errstate(divide="ignore", invalid="ignore"): + a = np.arctan((y / x).astype(np.float32)).astype(np.float32) + alt = np.where(y > 0, a + PI_F32, a - PI_F32).astype(np.float32) + return np.where(x < 0, alt, a).astype(np.float32) + + +def _heading_pos(xy: np.ndarray, y_col: np.ndarray, x_col: np.ndarray) -> np.ndarray: + """``[x, y, cos(h), sin(h)]`` with ``h = atan2_onnx(y_col, x_col)`` (float32).""" + h = atan2_onnx(y_col, x_col) + with np.errstate(invalid="ignore"): + return np.stack([xy[..., 0], xy[..., 1], np.cos(h), np.sin(h)], axis=-1).astype(np.float32) + + +def _onehot_pos(pos4: np.ndarray, cls: int) -> np.ndarray: + """Append the 10-way class one-hot (``add_class_type``).""" + onehot = np.zeros(pos4.shape[:-1] + (C.POS_CLASS_NUM,), np.float32) + onehot[..., cls] = 1.0 + return np.concatenate([pos4.astype(np.float32), onehot], axis=-1) + + +def line_features(x: np.ndarray) -> np.ndarray: + """``LineEncoder``: ``[points..., dx, dy]`` with dx / dy = next point minus point (0 at the last point).""" + x = np.asarray(x, np.float32) + d = np.zeros(x.shape[:-1] + (2,), np.float32) + d[..., :-1, 0] = x[..., 1:, 0] - x[..., :-1, 0] + d[..., :-1, 1] = x[..., 1:, 1] - x[..., :-1, 1] + return np.concatenate([x, d], axis=-1) + + +@dataclass +class EncoderFeatures: + """Host-side inputs of the encoder network (batch 1, no batch dim). ``valid[cat]`` is True for valid entities; + ``token_valid`` (564) gates the positional embedding, ``key_valid`` (564) is the fusion key mask (ego forced + valid); ``pos`` (564 x 14) are the position features, with invalid rows set to 0 (the graph's NaN rows there are + discarded by a ``Where``).""" + + ego: np.ndarray # [31, 4] truncated ego history + neighbor: np.ndarray # [320, 31, 9] + neighbor_type: np.ndarray # [320, 3] + static: np.ndarray # [5, 10] + lane: np.ndarray # [140, 20, 8] + lane_attr: np.ndarray # [140, 25] + lane_speed: np.ndarray # [140, 1] + lane_has_speed: np.ndarray # [140, 1] bool + route: np.ndarray # [25, 20, 8] + route_attr: np.ndarray # [25, 25] + route_speed: np.ndarray # [25, 1] + route_has_speed: np.ndarray # [25, 1] bool + polygon: np.ndarray # [10, 40, 5] + line_string: np.ndarray # [60, 20, 6] + goal: np.ndarray # [4] + ego_shape: np.ndarray # [3] + turn: np.ndarray # [30] + valid: Dict[str, np.ndarray] + token_valid: np.ndarray # [564] bool + key_valid: np.ndarray # [564] bool + pos: np.ndarray # [564, 14] + + def counts(self) -> Dict[str, int]: + return {k: int(v.sum()) for k, v in self.valid.items()} + + +@dataclass +class DecoderMasks: + agent_valid: np.ndarray # [321] bool: ego + neighbours with a non-zero current state (self-attention keys) + current_states: np.ndarray # [321, 4] normalised (prefix constraint) + + +def _any_nonzero(a: np.ndarray, axes) -> np.ndarray: + return np.any(a != 0, axis=axes) + + +def encoder_features(norm: Mapping[str, np.ndarray], masks: Mapping[str, np.ndarray]) -> EncoderFeatures: + """From the normalised inputs (``normalize_inputs``) and the speed masks (``speed_masks``), batch 1.""" + f32 = lambda k: np.asarray(norm[k], np.float32)[0] # noqa: E731 (drop the batch dim) + # ego: keep the 6 oldest rows (encoder.py:170-175); the token is always valid, its position row is all zero + ego = np.zeros((C.INPUT_T + 1, C.POSE_DIM), np.float32) + ego[C.EGO_HISTORY_KEEP] = f32("ego_agent_past")[C.EGO_HISTORY_KEEP] + pos = {"ego": _onehot_pos(ego[-1:], C.POS_CLASS["ego"])} + + # neighbours: keep the 6 newest rows (encoder.py:176-181, 441-451) + nb_raw = f32("neighbor_agents_past") + nb = np.zeros_like(nb_raw) + nb[:, C.NEIGHBOR_HISTORY_KEEP] = nb_raw[:, C.NEIGHBOR_HISTORY_KEEP] + nb_type = nb[:, -1, 8:11].copy() + x8 = nb[..., :8] + step_valid = _any_nonzero(x8, -1) # [320, 31] + nb_valid = step_valid.any(axis=-1) # [320] + feat = np.concatenate([x8, step_valid[..., None].astype(np.float32)], axis=-1) + feat[..., 4:6] = 0.0 # velocities zeroed after the validity test + feat[~nb_valid] = 0.0 + pos["neighbor"] = _onehot_pos(x8[:, -1, :4], C.POS_CLASS["neighbor"]) + + static = f32("static_objects") + st_valid = _any_nonzero(static[..., :10], -1) + static = np.where(st_valid[:, None], static, 0.0).astype(np.float32) + pos["static"] = _onehot_pos(f32("static_objects")[:, :4], C.POS_CLASS["static"]) + + def lanes(key: str, speed_key: str, mask_key: str, cat: str): + x = f32(key) + attr = x[:, 0, C.LANE_FEATURE_DIM:].copy() + x8 = x[..., :C.LANE_FEATURE_DIM] + valid = _any_nonzero(x8, (-1, -2)) + p = x8[:, C.LANE_POS_INDEX, :4] + pos[cat] = _onehot_pos(_heading_pos(p, p[:, 3], p[:, 2]), C.POS_CLASS[cat]) + x8 = np.where(valid[:, None, None], x8, 0.0).astype(np.float32) + speed = f32(speed_key).reshape(-1, 1) + has = np.asarray(masks[mask_key])[0].reshape(-1, 1).astype(bool) + return x8, attr, speed, has, valid + + lane, lane_attr, lane_speed, lane_has, lane_valid = lanes("lanes", "lanes_speed_limit", + "lanes_has_speed_limit", "lane") + route, route_attr, route_speed, route_has, route_valid = lanes("route_lanes", "route_lanes_speed_limit", + "route_lanes_has_speed_limit", "route") + + def lines(key: str, cat: str, pos_index: int): + x = line_features(f32(key)) + valid = _any_nonzero(x[..., :4], (-1, -2)) + p = x[:, pos_index, :4] + pos[cat] = _onehot_pos(_heading_pos(p, p[:, 3], p[:, 2]), C.POS_CLASS[cat]) + return np.where(valid[:, None, None], x, 0.0).astype(np.float32), valid + + polygon, poly_valid = lines("polygons", "polygon", C.POLYGON_POS_INDEX) + line_string, ls_valid = lines("line_strings", "line_string", C.LINE_STRING_POS_INDEX) + + goal = f32("goal_pose").reshape(-1) + pos["goal"] = _onehot_pos(goal[None], C.POS_CLASS["goal"]) + unit = np.array([[0.0, 0.0, 1.0, 0.0]], np.float32) + pos["ego_shape"] = _onehot_pos(unit, C.POS_CLASS["ego_shape"]) + pos["turn"] = _onehot_pos(unit, C.POS_CLASS["turn"]) + turn = np.asarray(norm["turn_indicators"], np.float32)[0, :C.TURN_INDICATOR_HISTORY].copy() + + one = np.ones(1, bool) + valid = {"ego": one, "neighbor": nb_valid, "static": st_valid, "lane": lane_valid, "route": route_valid, + "polygon": poly_valid, "line_string": ls_valid, "goal": one, "ego_shape": one.copy(), "turn": one.copy()} + token_valid = np.concatenate([valid[name] for name, _ in C.TOKEN_LAYOUT]) + key_valid = token_valid.copy() + key_valid[0] = True + pos_all = np.concatenate([pos[name] for name, _ in C.TOKEN_LAYOUT], axis=0).astype(np.float32) + pos_all[~token_valid] = 0.0 + return EncoderFeatures(ego, feat.astype(np.float32), nb_type, static, lane, lane_attr, lane_speed, lane_has, + route, route_attr, route_speed, route_has, polygon, line_string, goal, + f32("ego_shape").reshape(-1).copy(), turn, valid, token_valid, key_valid, pos_all) + + +def decoder_masks(norm: Mapping[str, np.ndarray]) -> DecoderMasks: + """Agent key mask of the DiT self-attention (``dit.py:155-156``; the ego is always valid) and the current states + of the prefix constraint (``multi_step_inference.cpp:300-325``), from the normalised inputs.""" + nb_now = np.asarray(norm["neighbor_agents_past"], np.float32)[0, :, C.INPUT_T, :C.POSE_DIM] + agent_valid = np.concatenate([np.ones(1, bool), np.any(nb_now != 0, axis=-1)]) + cs = np.zeros((C.MAX_NUM_AGENTS, C.POSE_DIM), np.float32) + cs[0] = np.asarray(norm["ego_current_state"], np.float32)[0, :C.POSE_DIM] + cs[1:] = nb_now + return DecoderMasks(agent_valid, cs) diff --git a/code/tt_diffusion_planner/host/normalize.py b/code/tt_diffusion_planner/host/normalize.py new file mode 100644 index 0000000000000000000000000000000000000000..1d841dea0ae05f4e047a4f8ed362ea5e79fadc36 --- /dev/null +++ b/code/tt_diffusion_planner/host/normalize.py @@ -0,0 +1,67 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Input normalization of the Autoware node (``preprocessing_utils.cpp:34-84``) and the speed-limit masks. + +The rule cannot be folded into the first-layer weights (SPEC 9.1, PLAN 2.12): each row of the last dimension is +normalised as ``(v - mean) / std`` in float32 unless every value of the row satisfies ``|v| < FLT_EPSILON``, in which +case the row is left untouched (padding stays zero while ``x`` has mean 10 m). ``ego_shape``, +``sampled_trajectories``, ``turn_indicators`` and ``delay`` are never normalised. +""" +from __future__ import annotations + +from typing import Dict, Mapping, Tuple + +import numpy as np + +from ..reference import config as C + +__all__ = ["FLT_EPSILON", "normalize_inputs", "normalize_array", "speed_masks"] + +FLT_EPSILON = np.float32(np.finfo(np.float32).eps) # std::numeric_limits::epsilon() + + +def normalize_array(value: np.ndarray, mean: np.ndarray, std: np.ndarray) -> np.ndarray: + """``normalize_vector`` of ``preprocessing_utils.cpp:36-63`` on one tensor (a new float32 array). + + Rows have ``std.size`` columns; a single-value mean / std broadcasts over the columns (the C++ ``mean.size() == 1`` + branch); a zero standard deviation raises like the C++.""" + data = np.array(value, dtype=np.float32, copy=True) + mean = np.asarray(mean, np.float32).reshape(-1) + std = np.asarray(std, np.float32).reshape(-1) + if mean.size != std.size: + raise ValueError("Mean and std must be same size") + cols = std.size + if cols == 0 or data.size % cols: + raise ValueError(f"data size {data.size} is not divisible by the normalizer size {cols}") + if np.any(np.abs(std) < FLT_EPSILON): + raise ValueError("Standard deviation is zero, cannot normalize data") + rows = data.reshape(-1, cols) + zero_row = np.all(np.abs(rows) < FLT_EPSILON, axis=1) + m = np.broadcast_to(mean if mean.size > 1 else mean[:1], (cols,)) + s = np.broadcast_to(std if std.size > 1 else std[:1], (cols,)) + normed = ((rows - m) / s).astype(np.float32) # float32 subtract then divide, element by element + rows[~zero_row] = normed[~zero_row] + return rows.reshape(data.shape) + + +def normalize_inputs(raw: Mapping[str, np.ndarray], + observation: Mapping[str, Tuple[np.ndarray, np.ndarray]]) -> Dict[str, np.ndarray]: + """``normalize_input_data(input_data_map, normalization_map)``: every key except the four skipped ones must have + a normalizer (``Missing key ... from normalization map`` otherwise).""" + out: Dict[str, np.ndarray] = {} + for key, value in raw.items(): + if key in C.SKIP_NORMALIZATION: + out[key] = np.array(value, dtype=np.float32, copy=True) + continue + if key not in observation: + raise KeyError(f"Missing key {key} from normalization map") + mean, std = observation[key] + out[key] = normalize_array(value, mean, std) + return out + + +def speed_masks(norm_inputs: Mapping[str, np.ndarray]) -> Dict[str, np.ndarray]: + """``lanes_has_speed_limit`` / ``route_lanes_has_speed_limit`` of the TensorRT path: normalised speed limit + ``> FLT_EPSILON`` (``inference/utils.hpp:112-123``; the ORT backend uses ``> 0``, + ``onnxruntime_inference.cpp:41-48``; Autoware's default backend is TensorRT).""" + return {"lanes_has_speed_limit": np.asarray(norm_inputs["lanes_speed_limit"]) > FLT_EPSILON, + "route_lanes_has_speed_limit": np.asarray(norm_inputs["route_lanes_speed_limit"]) > FLT_EPSILON} diff --git a/code/tt_diffusion_planner/host/pipeline.py b/code/tt_diffusion_planner/host/pipeline.py new file mode 100644 index 0000000000000000000000000000000000000000..8a939320e00699dadd5509a80db1a2de63a315e1 --- /dev/null +++ b/code/tt_diffusion_planner/host/pipeline.py @@ -0,0 +1,107 @@ +# SPDX-License-Identifier: Apache-2.0 +"""The host halves of one plan, shared by the Python API, the HTTP server and the CPU reference. + +``prepare(raw, normalization)`` -> :class:`Prepared`: the raw ONNX-named tensors (already checked against +``config.INPUT_SCHEMA``) normalised like the node, the speed masks, the encoder's host features, the decoder masks +and the solver's initial ``x``. The network (device trace or CPU reference) turns it into ``final_x0`` (normalised +``[321, 81, 4]``, prefix constraint applied) and the turn-indicator logits; ``make_output(...)`` applies the node's +post-processing and returns the ``ttaw.outputs.Trajectory`` that ``model(...)`` and ``POST /predict`` return. +""" +from __future__ import annotations + +from dataclasses import dataclass +from typing import Any, Dict, Mapping, Optional, Sequence + +import numpy as np + +from ..reference import config as C +from ..ttaw.io import InputError, encode_array +from ..ttaw.outputs import Trajectory +from .features import DecoderMasks, EncoderFeatures, decoder_masks, encoder_features +from .normalize import normalize_inputs, speed_masks +from .postprocess import (TurnIndicatorManager, denoising_steps_ego, denormalize, predicted_paths, + trajectory_from_poses) + +__all__ = ["Prepared", "prepare", "make_output", "TRAJECTORY_COLUMNS", "PREDICTED_AGENT_COLUMNS", "RUNTIME_PARAMS"] + +TRAJECTORY_COLUMNS = ("x", "y", "yaw", "cos", "sin", "velocity", "acceleration") +PREDICTED_AGENT_COLUMNS = ("x", "y", "yaw", "cos", "sin") +# RT-host parameters (request params / call kwargs): name -> (type, min, max, default); the node's YAML defaults +RUNTIME_PARAMS = { + "velocity_smoothing_window": (int, 1, C.OUTPUT_T - 1, C.VELOCITY_SMOOTHING_WINDOW), + "stopping_threshold": (float, 0.0, None, C.STOPPING_THRESHOLD), + "turn_indicator_keep_offset": (float, None, None, C.TURN_INDICATOR_KEEP_OFFSET), + "return_denoising_steps": (bool, None, None, False), +} + + +@dataclass +class Prepared: + """Host state of one plan.""" + + raw: Dict[str, np.ndarray] # the validated raw tensors (batch 1) + norm: Dict[str, np.ndarray] # normalised tensors + the two speed masks + features: EncoderFeatures + decoder: DecoderMasks + x_T: np.ndarray # [321, 81, 4] initial solver state (normalised; correction not yet applied) + neighbor_rows: np.ndarray # indices of the non-empty neighbour rows (the node's emitted agents) + enable_force_stop: bool # ego vx > DBL_EPSILON (diffusion_planner_core.cpp:628-629) + prev_report: int # the latest TurnIndicatorsReport (turn_indicators[30]) + + +def prepare(raw: Mapping[str, np.ndarray], observation: Mapping[str, Any]) -> Prepared: + """``raw``: the 15 tensors of ``config.INPUT_SCHEMA`` (float32, batch 1, ego frame, before normalization); + ``observation``: ``Normalization.observation``.""" + missing = sorted(set(C.INPUT_NAMES) - set(raw)) + if missing: + raise InputError(f"inputs: missing {missing}") + raw = {k: np.ascontiguousarray(np.asarray(raw[k], np.float32)) for k in C.INPUT_NAMES} + for k, shape in C.INPUT_SHAPES.items(): + if raw[k].shape != shape: + raise InputError(f"inputs[{k!r}] has shape {raw[k].shape}, expected {shape}") + try: + norm = normalize_inputs(raw, observation) + except (KeyError, ValueError) as e: # a normalizer problem is a weights-file problem, not a client mistake + raise RuntimeError(f"normalization failed: {e}") from e + norm.update(speed_masks(norm)) + feats = encoder_features(norm, norm) + dec = decoder_masks(norm) + nb = raw["neighbor_agents_past"][0].reshape(C.MAX_NUM_NEIGHBORS, -1) + rows = np.flatnonzero(np.any(nb != 0, axis=1)) + vx = float(raw["ego_current_state"][0, 4]) + prev_report = int(round(float(raw["turn_indicators"][0, C.INPUT_T]))) + x_T = norm["sampled_trajectories"][0].copy() + return Prepared(raw, norm, feats, dec, x_T, rows, vx > np.finfo(np.float64).eps, prev_report) + + +def make_output(final_x0: np.ndarray, logit: np.ndarray, prepared: Prepared, normalization: Any, + params: Mapping[str, Any], *, model: str = "", denoising_steps: Optional[Sequence[np.ndarray]] = None, + timing_ms: Optional[Dict[str, float]] = None, meta: Optional[Dict[str, Any]] = None) -> Trajectory: + """The node's post-processing of one plan -> ``Trajectory`` (base_link). + + ``final_x0``: ``[321, 81, 4]`` normalised (t = 0 = current state); ``logit``: ``[5]``; ``params``: the validated + ``RUNTIME_PARAMS``; ``denoising_steps``: the 11 iterates ``[321, 81, 4]`` when ``return_denoising_steps``.""" + x0 = np.asarray(final_x0, np.float32).reshape(C.MAX_NUM_AGENTS, C.OUTPUT_T + 1, C.POSE_DIM) + mean, std = normalization.state() + denorm = denormalize(x0, mean, std) # [321, 80, 4] + traj = trajectory_from_poses(denorm[0], (0.0, 0.0, 0.0), + velocity_smoothing_window=int(params["velocity_smoothing_window"]), + enable_force_stop=prepared.enable_force_stop, + stopping_threshold=float(params["stopping_threshold"])) + manager = TurnIndicatorManager(keep_offset=float(params["turn_indicator_keep_offset"])) + decision = manager.evaluate(np.asarray(logit, np.float32).reshape(-1), 0.0, prepared.prev_report) + rows = prepared.neighbor_rows + paths = predicted_paths(denorm, C.MAX_NUM_NEIGHBORS)[rows] if rows.size else np.zeros( + (0, C.OUTPUT_T, len(PREDICTED_AGENT_COLUMNS)), np.float32) + out_meta: Dict[str, Any] = {"predicted_agent_columns": list(PREDICTED_AGENT_COLUMNS), + "predicted_agent_rows": [int(r) for r in rows], + "force_stop": bool(traj.force_stop), + "time_from_start_s": [round(float(t), 3) for t in traj.time_from_start], + "valid_counts": prepared.features.counts()} + if params.get("return_denoising_steps") and denoising_steps is not None: + ego_steps = denoising_steps_ego(np.stack(denoising_steps), mean, std) + out_meta["denoising_steps"] = encode_array(ego_steps, key="denoising_steps") + out_meta.update(meta or {}) + return Trajectory(traj.as_columns(), columns=TRAJECTORY_COLUMNS, turn_indicator=decision.to_dict(), + predicted_agents=paths, model=model, frame_id="base_link", timing_ms=dict(timing_ms or {}), + meta=out_meta) diff --git a/code/tt_diffusion_planner/host/postprocess.py b/code/tt_diffusion_planner/host/postprocess.py new file mode 100644 index 0000000000000000000000000000000000000000..bd77f571b35a6178a3e41878400f4274526b7026 --- /dev/null +++ b/code/tt_diffusion_planner/host/postprocess.py @@ -0,0 +1,222 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Post-processing of the Autoware node (``postprocessing_utils.cpp``, ``turn_indicator_manager.cpp``, +``diffusion_planner_core.cpp:605-694``), in the ego frame of the input tensors. + +The node transforms every predicted pose to the map frame before building messages; the bundle has no map pose, so +it applies the same rules with an identity ego-to-map transform (``base_link`` output). Distances, velocities and +accelerations are invariant to that rigid transform; the published orientation is not exactly (see below). + +Rules reproduced, with their numeric types: + +- denormalisation: drop t = 0 (unless kept), ``v * std + mean`` per agent and pose dim, float32; +- poses: 4x4 matrices in double from the RAW (unnormalised) network cos / sin; the quaternion is Eigen's conversion + of that 3x3 block, which does not normalise, so for ``|(cos, sin)| != 1`` (0.84-1.09 in the goldens, SPEC 4.6.6) + the heading a consumer reads with ``tf2::getYaw`` differs from ``atan2(sin, cos)``; ``yaw`` here is that + ``tf2::getYaw`` value of the identity-transform quaternion; +- ego trajectory (``get_trajectory_from_poses``): ``time_from_start = 0.1 (i + 1)``; velocity = 3-D distance to the + previous point (the first point to the ego position) / 0.1, stored as float; forward moving average over + ``velocity_smoothing_window`` points summed in double; force stop (only when the ego moves: ``vx > DBL_EPSILON``) + once the smoothed velocity crosses below ``stopping_threshold``: velocity 0 and the pose frozen to the previous + pose; the last ``window - 1`` points keep the last smoothed velocity; acceleration = forward difference / 0.1 in + double stored as float, 0 for the last point; +- predicted objects: one 80-point path per emitted neighbour (the first ``n`` rows of the neighbour tensor, sorted + by distance), window 1, no force stop (only the poses are published); +- turn indicator (``TurnIndicatorManager::evaluate``): hold the last non-KEEP command for ``hold_duration`` (state + across calls), else add ``keep_offset`` to the KEEP logit, ``p_i = expf(l_i - max)``, ``p_i /= (1e-4f + sum)``, + argmax (first maximum); KEEP repeats the previous TurnIndicatorsReport, otherwise the command is the index; +- ``~/debug/denoising_steps``: the ego row of every solver iterate, denormalised with t = 0 kept. +""" +from __future__ import annotations + +import math +from dataclasses import dataclass, field +from typing import Dict, Optional, Tuple + +import numpy as np + +from ..reference import config as C +from .solver import SCALAR_MATH + +__all__ = ["denormalize", "quaternion_from_cos_sin", "tf2_yaw", "EgoTrajectory", "trajectory_from_poses", + "predicted_paths", "TurnIndicatorManager", "TurnDecision", "denoising_steps_ego"] + +DBL_EPSILON = float(np.finfo(np.float64).eps) + + +def denormalize(x: np.ndarray, state_mean: np.ndarray, state_std: np.ndarray, *, + keep_current_state: bool = False) -> np.ndarray: + """``denormalize_prediction``: ``x`` ``[321, 81, 4]`` (normalised, t = 0 = current state) -> ``[321, 80, 4]`` + metres (``[321, 81, 4]`` with ``keep_current_state``). ``state_mean`` / ``state_std`` broadcast as + ``[321 or 1, 1, 4]``.""" + x = np.asarray(x, np.float32) + if x.shape[-3:] != (C.MAX_NUM_AGENTS, C.OUTPUT_T + 1, C.POSE_DIM): + raise ValueError(f"unsupported prediction shape {x.shape}") + p = x if keep_current_state else x[..., 1:, :] + return (p * np.asarray(state_std, np.float32) + np.asarray(state_mean, np.float32)).astype(np.float32) + + +def quaternion_from_cos_sin(cos_yaw: np.ndarray, sin_yaw: np.ndarray) -> np.ndarray: + """Eigen ``Quaterniond(rotation_matrix)`` of ``[[c, -s, 0], [s, c, 0], [0, 0, 1]]`` (no normalisation), in double: + ``[..., 4]`` = (x, y, z, w). Eigen's branch: trace = 2c + 1 > 0 -> w = sqrt(trace + 1) / 2, z = (m10 - m01) / (4 w); + else the largest diagonal is m22 -> z = sqrt(m22 - m00 - m11 + 1) / 2, w = (m10 - m01) / (4 z).""" + c = np.asarray(cos_yaw, np.float64) + s = np.asarray(sin_yaw, np.float64) + q = np.zeros(c.shape + (4,), np.float64) + trace = c + c + 1.0 + pos = trace > 0.0 + with np.errstate(divide="ignore", invalid="ignore"): + t = np.sqrt(trace + 1.0) + w_pos, z_pos = 0.5 * t, (s - (-s)) * (0.5 / t) + # trace <= 0: i = 2 (m22 = 1 > m00 = c, since c <= -0.5), j = 0, k = 1 + t2 = np.sqrt(1.0 - c - c + 1.0) + z_neg, w_neg = 0.5 * t2, (s - (-s)) * (0.5 / t2) + q[..., 2] = np.where(pos, z_pos, z_neg) + q[..., 3] = np.where(pos, w_pos, w_neg) + return q + + +def tf2_yaw(q: np.ndarray) -> np.ndarray: + """``tf2::getYaw`` (scale-invariant): ``atan2(2 (x y + w z), w^2 + x^2 - y^2 - z^2)``.""" + x, y, z, w = (q[..., i] for i in range(4)) + return np.arctan2(2.0 * (x * y + w * z), w * w + x * x - y * y - z * z) + + +@dataclass +class EgoTrajectory: + """One trajectory as the node publishes it (in the ego frame here): positions [N, 3] (double), the raw cos / sin + of each pose [N, 2] (after force-stop freezing), quaternions [N, 4] (x, y, z, w), yaw [N] (``tf2::getYaw``), + ``velocity`` / ``acceleration`` [N] float32, ``time_from_start`` [N] seconds, and whether force stop fired.""" + + position: np.ndarray + cos_sin: np.ndarray + quaternion: np.ndarray + yaw: np.ndarray + velocity: np.ndarray + acceleration: np.ndarray + time_from_start: np.ndarray + force_stop: bool + + def as_columns(self) -> np.ndarray: + """``[N, 7]`` float32: x, y, yaw, cos, sin, velocity, acceleration (the ``Trajectory.poses`` layout).""" + return np.stack([self.position[:, 0], self.position[:, 1], self.yaw, self.cos_sin[:, 0], self.cos_sin[:, 1], + self.velocity, self.acceleration], axis=-1).astype(np.float32) + + +def trajectory_from_poses(poses_xycs: np.ndarray, base_position: Tuple[float, float, float] = (0.0, 0.0, 0.0), *, + velocity_smoothing_window: int = C.VELOCITY_SMOOTHING_WINDOW, + enable_force_stop: bool = True, + stopping_threshold: float = C.STOPPING_THRESHOLD) -> EgoTrajectory: + """``get_trajectory_from_poses`` (``postprocessing_utils.cpp:362-455``) for poses ``[N, 4]`` = denormalised + (x, y, cos, sin) float32 of one agent, positions at z = ``base_position[2]`` (identity transform).""" + p = np.asarray(poses_xycs, np.float32) + n = p.shape[0] + dt = C.TRAJECTORY_DT + pos = np.zeros((n, 3), np.float64) + pos[:, 0] = p[:, 0].astype(np.float64) + pos[:, 1] = p[:, 1].astype(np.float64) + pos[:, 2] = float(base_position[2]) + cs = p[:, 2:4].astype(np.float64).copy() + quat = quaternion_from_cos_sin(cs[:, 0], cs[:, 1]) + vel = np.zeros(n, np.float32) + prev = np.asarray(base_position, np.float64) + for i in range(n): + d = math.hypot(pos[i, 0] - prev[0], pos[i, 1] - prev[1], pos[i, 2] - prev[2]) + vel[i] = np.float32(d / dt) + prev = pos[i] + w = int(velocity_smoothing_window) + if n <= w: + raise ValueError("velocity_smoothing_window must be smaller than number of points") + thr = np.float32(stopping_threshold) + force_stop = False + for i in range(0, n - w + 1): + acc = 0.0 + for k in range(w): + acc += float(vel[i + k]) + vel[i] = np.float32(acc / float(w)) + if enable_force_stop and i > 0 and abs(vel[i - 1]) > thr and abs(vel[i]) < thr: + force_stop = True + if i > 0 and force_stop: + vel[i] = np.float32(0.0) + pos[i], cs[i], quat[i] = pos[i - 1], cs[i - 1], quat[i - 1] + last = vel[n - w] + for i in range(n - w + 1, n): + vel[i] = last + if force_stop: + vel[i] = np.float32(0.0) + pos[i], cs[i], quat[i] = pos[i - 1], cs[i - 1], quat[i - 1] + accel = np.zeros(n, np.float32) + for i in range(n - 1): + accel[i] = np.float32((float(vel[i + 1]) - float(vel[i])) / dt) + tfs = dt * (np.arange(n, dtype=np.float64) + 1.0) + return EgoTrajectory(pos, cs, quat, tf2_yaw(quat), vel, accel, tfs, force_stop) + + +def predicted_paths(denorm: np.ndarray, n_neighbors: int) -> np.ndarray: + """``create_predicted_objects`` poses for the first ``n_neighbors`` neighbours: ``[n, 80, 5]`` float32 = + x, y, yaw (``tf2::getYaw``), cos, sin. ``denorm`` is ``[321, 80, 4]`` (agent 0 = ego).""" + nb = np.asarray(denorm, np.float32)[1:1 + int(n_neighbors)] + q = quaternion_from_cos_sin(nb[..., 2], nb[..., 3]) + return np.stack([nb[..., 0], nb[..., 1], tf2_yaw(q).astype(np.float32), nb[..., 2], nb[..., 3]], + axis=-1).astype(np.float32) + + +def denoising_steps_ego(iterates: np.ndarray, state_mean: np.ndarray, state_std: np.ndarray) -> np.ndarray: + """``create_denoising_steps_message`` (batch 1): the ego row of each iterate, denormalised with t = 0 kept: + ``[steps, 81, 4]``. ``iterates`` is ``[steps, 321, 81, 4]`` or just the ego rows ``[steps, 1, 81, 4]`` (what the + device plan reads back); the float32 arithmetic is the same element by element (``x * std + mean``).""" + it = np.asarray(iterates, np.float32) + if it.ndim != 4 or it.shape[1] not in (1, C.MAX_NUM_AGENTS) or it.shape[2:] != (C.OUTPUT_T + 1, C.POSE_DIM): + raise ValueError(f"unsupported iterates shape {it.shape}") + shape = (C.MAX_NUM_AGENTS, 1, C.POSE_DIM) + mean0 = np.broadcast_to(np.asarray(state_mean, np.float32), shape)[0] + std0 = np.broadcast_to(np.asarray(state_std, np.float32), shape)[0] + return (it[:, 0] * std0 + mean0).astype(np.float32) + + +@dataclass +class TurnDecision: + command: int + logits: Tuple[float, ...] + probabilities: Tuple[float, ...] + keep_selected: bool + held: bool = False + + def to_dict(self) -> Dict[str, object]: + name = C.TURN_INDICATOR_COMMAND_NAMES.get(self.command, str(self.command)) + return {"command": int(self.command), "command_name": name, "keep_selected": bool(self.keep_selected), + "held": bool(self.held), "logits": [float(v) for v in self.logits], + "probabilities": [float(v) for v in self.probabilities]} + + +@dataclass +class TurnIndicatorManager: + """``TurnIndicatorManager`` (stateful: the last non-KEEP command and its stamp). A fresh manager has no held + command, which is what a single stateless request sees.""" + + hold_duration_s: float = C.TURN_INDICATOR_HOLD_DURATION_S + keep_offset: float = C.TURN_INDICATOR_KEEP_OFFSET + _last_command: int = field(default=0, repr=False) + _last_stamp_s: Optional[float] = field(default=None, repr=False) + + def evaluate(self, logit: np.ndarray, stamp_s: float = 0.0, + prev_report: int = C.TURN_INDICATORS_REPORT_DISABLE) -> TurnDecision: + lg = np.asarray(logit, np.float32).reshape(-1).copy() + raw = tuple(float(v) for v in lg) + if lg.size == 0: + return TurnDecision(1, raw, (), False) # TurnIndicatorsCommand::DISABLE + if self._last_stamp_s is not None and self._last_stamp_s > 0 and stamp_s <= self._last_stamp_s + float( + self.hold_duration_s): + return TurnDecision(self._last_command, raw, (), False, held=True) + lg[C.TURN_INDICATOR_OUTPUT_KEEP] = np.float32(lg[C.TURN_INDICATOR_OUTPUT_KEEP] + np.float32(self.keep_offset)) + mx = lg.max() + prob = np.array([SCALAR_MATH.exp(np.float32(v - mx)) for v in lg], np.float32) + total = np.float32(0.0001) + for v in prob: + total = np.float32(total + v) + prob = (prob / total).astype(np.float32) + idx = int(np.argmax(prob)) # std::max_element: the first maximum + keep = idx == C.TURN_INDICATOR_OUTPUT_KEEP + command = (int(prev_report) & 0xFF) if keep else idx + if not keep: + self._last_command, self._last_stamp_s = command, float(stamp_s) + return TurnDecision(command, raw, tuple(float(v) for v in prob), keep) diff --git a/code/tt_diffusion_planner/host/solver.py b/code/tt_diffusion_planner/host/solver.py new file mode 100644 index 0000000000000000000000000000000000000000..701ea8e04ef40b424bfa20604652f4186ad2a2ad --- /dev/null +++ b/code/tt_diffusion_planner/host/solver.py @@ -0,0 +1,273 @@ +# SPDX-License-Identifier: Apache-2.0 +"""DPM-Solver++(2M) with denoise-to-zero, as Autoware runs it around the decoder (multi_step mode). + +Port of ``PKG/src/inference/solver/dpm_solver.cpp`` (guidance disabled, Autoware's default and decision D11) and of +the prefix constraint of ``PKG/src/inference/multi_step_inference.cpp:300-340``. The scalar schedule is float32 math +evaluated with the C library's ``expf`` / ``logf`` / ``expm1f`` / ``sqrtf`` (what ``std::exp(float)`` etc. call in +the node) in the C++ operation order, so the timesteps and update coefficients are the node's bit for bit; numpy's +own float32 transcendentals differ from glibc in the last bit for many arguments. Without a loadable libm the module +falls back to numpy (``SCALAR_MATH`` says which). + +For ``steps = 10``: 11 decoder evaluations at ``timesteps[0..9]`` and ``1/N = 0.001`` (denoise-to-zero); the 11 +published iterates (``denoising_steps``) are the corrected ``x`` after each update and after the final evaluation. +The update formulas, with ``x`` and the model outputs float32 arrays and every scalar a float32:: + + first (step 1): x = (sigma_t / sigma_s) * x - (alpha_t * phi_1) * m + second (steps 2..): d = (m0 - m1) / r0 + x = (sigma_t / sigma_0) * x - (alpha_t * phi_1) * m0 - (0.5 * (alpha_t * phi_1)) * d + +:class:`SolverPlan` lists these scalars per update (the RT-dev coefficient tables of the on-device loop). +""" +from __future__ import annotations + +import ctypes +import ctypes.util +from dataclasses import dataclass, field +from typing import Callable, List, Optional, Tuple + +import numpy as np + +from ..reference import config as C + +__all__ = ["SCALAR_MATH", "marginal_log_mean_coeff", "marginal_alpha", "marginal_std", "marginal_lambda", + "inverse_lambda", "log_snr_timesteps", "SolverUpdate", "SolverPlan", "solver_plan", "first_update", + "second_update", "dpm_solver_sample", "apply_prefix_constraint", "SampleResult"] + +f32 = np.float32 +_T, _N = C.NOISE_SCHEDULE_T, C.NOISE_SCHEDULE_TOTAL_N +_B0, _B1 = C.NOISE_SCHEDULE_BETA0, C.NOISE_SCHEDULE_BETA1 + + +class _ScalarMath: + """float32 ``exp`` / ``log`` / ``expm1`` / ``sqrt`` of the C library, loaded on first use (numpy fallback).""" + + def __init__(self) -> None: + self._fns = None + self.source = "unloaded" + + def _load(self) -> None: + fns = None + try: + name = ctypes.util.find_library("m") or "libm.so.6" + libm = ctypes.CDLL(name) + fns = {} + for fn in ("expf", "logf", "expm1f", "sqrtf"): + f = getattr(libm, fn) + f.restype, f.argtypes = ctypes.c_float, [ctypes.c_float] + fns[fn] = f + self.source = f"libm ({name})" + except (OSError, AttributeError): + fns = None + self.source = "numpy (no C math library found)" + self._fns = fns or {} + + def _call(self, fn: str, x: float, fallback: Callable) -> np.float32: + if self._fns is None: + self._load() + if fn in self._fns: + return f32(self._fns[fn](float(f32(x)))) + return f32(fallback(f32(x))) + + def exp(self, x): + return self._call("expf", x, np.exp) + + def log(self, x): + return self._call("logf", x, np.log) + + def expm1(self, x): + return self._call("expm1f", x, np.expm1) + + def sqrt(self, x): + return self._call("sqrtf", x, np.sqrt) + + +SCALAR_MATH = _ScalarMath() +_m = SCALAR_MATH + + +# ---- noise schedule (dpm_solver.cpp:49-94), float32 in the C++ operation order --------------------------------------- +def marginal_log_mean_coeff(t) -> np.float32: + """``-0.25f * t * t * (beta1 - beta0) - 0.5f * t * beta0``.""" + t = f32(t) + return f32(f32(f32(f32(-0.25) * t) * t) * f32(_B1 - _B0)) - f32(f32(f32(0.5) * t) * _B0) + + +def marginal_alpha(t) -> np.float32: + return _m.exp(marginal_log_mean_coeff(t)) + + +def marginal_std(t) -> np.float32: + return _m.sqrt(f32(f32(1.0) - _m.exp(f32(f32(2.0) * marginal_log_mean_coeff(t))))) + + +def marginal_lambda(t) -> np.float32: + lmc = marginal_log_mean_coeff(t) + log_std = f32(f32(0.5) * _m.log(f32(f32(1.0) - _m.exp(f32(f32(2.0) * lmc))))) + return f32(lmc - log_std) + + +def _log_add_exp(a, b) -> np.float32: + m = max(f32(a), f32(b)) + return f32(m + _m.log(f32(_m.exp(f32(f32(a) - m)) + _m.exp(f32(f32(b) - m))))) + + +def inverse_lambda(lam) -> np.float32: + beta_delta = f32(_B1 - _B0) + tmp = f32(f32(f32(2.0) * beta_delta) * _log_add_exp(f32(f32(-2.0) * f32(lam)), f32(0.0))) + delta = f32(f32(_B0 * _B0) + tmp) + return f32(f32(tmp / f32(_m.sqrt(delta) + _B0)) / beta_delta) + + +def log_snr_timesteps(steps: int) -> List[np.float32]: + """``steps + 1`` times uniform in log-SNR between t = 1 and t = 1/N (``dpm_solver.cpp:80-94``).""" + t0 = f32(f32(1.0) / _N) + lam_t, lam_0 = marginal_lambda(_T), marginal_lambda(t0) + out = [] + for i in range(steps + 1): + ratio = f32(f32(i) / f32(steps)) + out.append(inverse_lambda(f32(lam_t + f32(f32(lam_0 - lam_t) * ratio)))) + return out + + +# ---- updates (dpm_solver.cpp:96-136) ------------------------------------------------------------------------------ +def first_update(x_s: np.ndarray, model_s: np.ndarray, s, t) -> np.ndarray: + h = f32(marginal_lambda(t) - marginal_lambda(s)) + sigma_s, sigma_t, alpha_t = marginal_std(s), marginal_std(t), marginal_alpha(t) + phi_1 = _m.expm1(f32(-h)) + a, b = f32(sigma_t / sigma_s), f32(alpha_t * phi_1) + return (a * x_s.astype(np.float32) - b * model_s.astype(np.float32)).astype(np.float32) + + +def second_update(x_s: np.ndarray, model_prev: Tuple[np.ndarray, np.ndarray], t_prev: Tuple[float, float], + t) -> np.ndarray: + m1, m0 = model_prev # model_prev_list[0] (older), model_prev_list[1] (newer) + t1, t0 = t_prev + lam1, lam0, lam_t = marginal_lambda(t1), marginal_lambda(t0), marginal_lambda(t) + sigma0, sigma_t, alpha_t = marginal_std(t0), marginal_std(t), marginal_alpha(t) + h0, h = f32(lam0 - lam1), f32(lam_t - lam0) + r0 = f32(h0 / h) + phi_1 = _m.expm1(f32(-h)) + a, b = f32(sigma_t / sigma0), f32(alpha_t * phi_1) + c = f32(f32(0.5) * b) + d1_0 = ((m0.astype(np.float32) - m1.astype(np.float32)) / r0).astype(np.float32) + return (a * x_s.astype(np.float32) - b * m0.astype(np.float32) - c * d1_0).astype(np.float32) + + +@dataclass(frozen=True) +class SolverUpdate: + """One solver update to time ``t``: ``x = a * x - b * m0 - c * (m0 - m1) / r0`` (``c = 0`` for the first-order + update, which has no ``m1``).""" + + order: int + t: float + a: float + b: float + c: float + r0: float + + +@dataclass(frozen=True) +class SolverPlan: + """Everything the on-device loop needs for ``steps`` (the RT-dev tables of PLAN.md 2.12): the decoder evaluation + times (``eval_times[k]`` feeds evaluation ``k``; the t-embedding / adaLN tables are built for these values), the + update after each of the first ``steps`` evaluations, and the published iterate times.""" + + steps: int + timesteps: Tuple[float, ...] + eval_times: Tuple[float, ...] + updates: Tuple[SolverUpdate, ...] + denoising_timesteps: Tuple[float, ...] + scalar_math: str = field(default="") + + +def solver_plan(steps: int = C.DPM_SOLVER_STEPS) -> SolverPlan: + if steps < C.DPM_SOLVER_ORDER: + raise ValueError("DpmSolver steps must be greater than or equal to solver order.") + ts = log_snr_timesteps(steps) + updates = [] + for step in range(1, steps + 1): + t = ts[step] + if step < C.DPM_SOLVER_ORDER: + s = _T if step == 1 else ts[step - 1] + h = f32(marginal_lambda(t) - marginal_lambda(s)) + a = f32(marginal_std(t) / marginal_std(s)) + b = f32(marginal_alpha(t) * _m.expm1(f32(-h))) + updates.append(SolverUpdate(1, float(t), float(a), float(b), 0.0, 1.0)) + else: + t1 = _T if step - 2 == 0 else ts[step - 2] + t0 = ts[step - 1] + lam1, lam0, lam_t = marginal_lambda(t1), marginal_lambda(t0), marginal_lambda(t) + h0, h = f32(lam0 - lam1), f32(lam_t - lam0) + b = f32(marginal_alpha(t) * _m.expm1(f32(-h))) + updates.append(SolverUpdate(2, float(t), float(f32(marginal_std(t) / marginal_std(t0))), float(b), + float(f32(f32(0.5) * b)), float(f32(h0 / h)))) + final_t = f32(f32(1.0) / _N) + return SolverPlan(steps, tuple(float(t) for t in ts), tuple(float(t) for t in ts[:steps]) + (float(final_t),), + tuple(updates), tuple(float(t) for t in ts[1:]) + (float(final_t),), SCALAR_MATH.source) + + +# ---- the loop (dpm_solver.cpp:138-234) --------------------------------------------------------------------------- +@dataclass +class SampleResult: + final_x: np.ndarray + denoising_steps: List[np.ndarray] + denoising_timesteps: List[float] + eval_times: List[float] + nfe: int + + +def apply_prefix_constraint(x: np.ndarray, current_states: np.ndarray) -> np.ndarray: + """``x[agent, 0, :] = current_states[agent]`` in place (``multi_step_inference.cpp:327-340``); ``x`` is + ``[321, 81, 4]`` (or with a leading batch dim).""" + x[..., 0, :] = current_states + return x + + +def dpm_solver_sample(initial_x: np.ndarray, model_fn: Callable[[np.ndarray, np.float32], np.ndarray], + correcting_fn: Callable[[np.ndarray], None], steps: int = C.DPM_SOLVER_STEPS, + on_iterate: Optional[Callable[[int, np.ndarray], None]] = None) -> SampleResult: + """``DpmSolver::sample`` without guidance. ``model_fn(x, t)`` is one decoder evaluation (x0 prediction), + ``correcting_fn(x)`` the in-place prefix constraint.""" + if steps < C.DPM_SOLVER_ORDER: + raise ValueError("DpmSolver steps must be greater than or equal to solver order.") + x = np.array(initial_x, dtype=np.float32, copy=True) + correcting_fn(x) + ts = log_snr_timesteps(steps) + steps_out: List[np.ndarray] = [] + step_ts: List[float] = [] + eval_times: List[float] = [] + + def evaluate(xx: np.ndarray, t) -> np.ndarray: + eval_times.append(float(t)) + return np.asarray(model_fn(xx, f32(t)), np.float32) + + def record(xx: np.ndarray, t) -> None: + steps_out.append(xx.copy()) + step_ts.append(float(t)) + if on_iterate is not None: + on_iterate(len(steps_out) - 1, xx) + + t_prev = [_T] + m_prev = [evaluate(x, ts[0])] + for step in range(1, C.DPM_SOLVER_ORDER): + t = ts[step] + x = first_update(x, m_prev[-1], t_prev[-1], t) + correcting_fn(x) + record(x, t) + t_prev.append(t) + m_prev.append(evaluate(x, t)) + for step in range(C.DPM_SOLVER_ORDER, steps + 1): + t = ts[step] + x = second_update(x, (m_prev[0], m_prev[1]), (t_prev[0], t_prev[1]), t) + correcting_fn(x) + record(x, t) + t_prev = [t_prev[1], t] + m_prev = [m_prev[1], None] + if step < steps: + m_prev[1] = evaluate(x, t) + final_t = f32(f32(1.0) / _N) + x = evaluate(x, final_t) # denoise to zero: the model output itself, no guidance on the last step + x = np.array(x, dtype=np.float32, copy=True) + correcting_fn(x) + record(x, final_t) + return SampleResult(x, steps_out, step_ts, eval_times, len(eval_times)) diff --git a/code/tt_diffusion_planner/io.py b/code/tt_diffusion_planner/io.py new file mode 100644 index 0000000000000000000000000000000000000000..15a19ba74d613fba96399f74edddd99babaa08aa --- /dev/null +++ b/code/tt_diffusion_planner/io.py @@ -0,0 +1,32 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Input decoding and output encoding shared by the Python API and the HTTP server of diffusion-planner-p150. + +The implementation is the vendored ``ttaw.io`` (C08); numpy only, data parsing only, and every client mistake raises +:class:`InputError`, which the server maps to HTTP 400. This model's only input is ``inputs``: the 15 ONNX-named +planner tensors of the node's ``create_input_data()`` (``reference.config.INPUT_SCHEMA``), as a ``{name: array}`` +mapping, an ``.npz`` path or its bytes, or the JSON envelope ``{"format": "npz", "data": }`` / +``{"format": "json", "arrays": {...}}``. :func:`load_inputs` / :func:`decode_inputs` bind the schema, so +``model(inputs=...)`` and ``POST /predict`` accept and refuse the same inputs (names, shapes, finite values). +""" +from __future__ import annotations + +from typing import Any, Mapping, Optional + +from .reference.config import INPUT_SCHEMA +from .ttaw import io as _io +from .ttaw.io import * # noqa: F401,F403 (the decoders and encoders: ttaw/API.md section 9) + +__all__ = list(_io.__all__) + ["DEFAULT_POINT_FIELDS", "INPUT_SCHEMA", "load_inputs", "decode_inputs"] + +DEFAULT_POINT_FIELDS = () # no point cloud input + + +def load_inputs(source: Any) -> dict: + """Python-API planner input (mapping, ``.npz`` path or bytes, JSON envelope) -> ``{name: float32 array}`` + checked against ``INPUT_SCHEMA``.""" + return _io.load_named_arrays(source, INPUT_SCHEMA) + + +def decode_inputs(spec: Mapping[str, Any], *, max_bytes: Optional[int] = None) -> dict: + """The ``inputs`` envelope of ``/predict`` -> ``{name: float32 array}`` checked against ``INPUT_SCHEMA``.""" + return _io.decode_named_arrays(spec, INPUT_SCHEMA, max_bytes=max_bytes) diff --git a/code/tt_diffusion_planner/reference/__init__.py b/code/tt_diffusion_planner/reference/__init__.py new file mode 100644 index 0000000000000000000000000000000000000000..be4a22d422298f64b00d0198e3165ee1f24225bb --- /dev/null +++ b/code/tt_diffusion_planner/reference/__init__.py @@ -0,0 +1,34 @@ +# SPDX-License-Identifier: Apache-2.0 +"""CPU reference of Diffusion Planner v5.0 (Autoware diffusion_planner) -- the ground truth every PCC / +output-agreement gate of the TT port compares against. Importable without ttnn (torch, numpy, onnx only). + +- ``config.py`` dimensions, token layout, constants and node parameters, each with its Autoware / ONNX source. +- ``weights.py`` the three ONNX files and ``diffusion_planner.param.json`` read as DATA with the vendored + ``ttaw.weights.OnnxWeights``, every tensor addressed through its consuming node, into one canonical + ``{name: float32}`` dict that the reference and the ttnn graph share (no BatchNorm to fold). +- ``model.py`` pure-PyTorch fp32 encoder / DiT decoder / turn head with per-module taps. +- ``rewrites.py`` the exact graph rewrites the TT port applies (per-step adaLN tables folded into the LayerNorm + affine, hoisted cross-attention K/V, the pad-relative fp32 pre-projection island) as float64-built + constants plus CPU forwards that use them, tested against ``model.py``. +- ``pipeline.py`` ``ReferencePlanner``: host pre-processing (``..host``) -> encoder -> DPM-Solver++(2M) loop over 11 + decoder evaluations -> turn head -> host post-processing; ``run()`` records taps for goldens. +- ``ort.py`` ONNX Runtime on the shipped ONNX (the reference's oracle; research venv and tests only). +- ``goldens.py`` golden generation (taps + outputs per scene) and the small goldens kept in ``tests/goldens``. + +The ``to_dict()`` of ``ReferencePlanner()(inputs=sample)`` is stored as ``samples/.reference.json``: +``server/smoke_test.py`` compares the served output with it. +""" + +__all__ = ["ReferencePlanner", "load_weights", "find_weights_dir"] + + +def __getattr__(name): + if name == "ReferencePlanner": + from .pipeline import ReferencePlanner + + return ReferencePlanner + if name in ("load_weights", "find_weights_dir"): + from . import weights + + return getattr(weights, name) + raise AttributeError(f"module {__name__!r} has no attribute {name!r}") diff --git a/code/tt_diffusion_planner/reference/config.py b/code/tt_diffusion_planner/reference/config.py new file mode 100644 index 0000000000000000000000000000000000000000..04f73581c6ba5309ab74db166cf72288d03b3685 --- /dev/null +++ b/code/tt_diffusion_planner/reference/config.py @@ -0,0 +1,175 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Constants of the Autoware Diffusion Planner v5.0 contract, each with its source. + +Shared by the host pre/post-processing (``tt_diffusion_planner.host``), the CPU reference (``reference``) and the +ttnn graph (``tt``). Every value that changes a device shape is a COMPILE parameter (PLAN.md section 0.4): a change +means a new trace and a new image. + +Source abbreviations: ``PKG`` = autoware_universe @ 9ceaccf ``planning/autoware_diffusion_planner``; ``HFD`` = the +weights ``AutowareFoundation/diffusion_planner@v5.0`` (423efde6); ``SPEC`` = ``research/diffusion-planner/SPEC.md`` of +the porting workspace; ``T4M`` = tier4/Diffusion-Planner @ 40114a8 ``diffusion_planner/diffusion_planner`` (read +for semantics only, never imported). +""" +from __future__ import annotations + +from typing import Dict, Tuple + +import numpy as np + +# ---- dimensions (PKG/include/autoware/diffusion_planner/dimensions.hpp:26-105) -------------------------------------- +NUM_SEGMENTS_IN_LANE = 140 +NUM_SEGMENTS_IN_ROUTE = 25 +NUM_POLYGONS = 10 +NUM_LINE_STRINGS = 60 +NUM_STATIC_OBJECTS = 5 +MAX_NUM_NEIGHBORS = 320 +MAX_NUM_AGENTS = MAX_NUM_NEIGHBORS + 1 # ego + neighbours +HIDDEN_DIM = 256 +POINTS_PER_SEGMENT = 20 +POINTS_PER_POLYGON = 40 +POINTS_PER_LINE_STRING = 20 +LINE_TYPE_NUM = 10 +POLYGON_TYPE_NUM = 1 # intersection_area +LINE_STRING_TYPE_NUM = 2 # stop_line, road_border +SEGMENT_POINT_DIM = 13 + 2 * LINE_TYPE_NUM # 33 +INPUT_T = 30 # history steps before the current one (31 samples with the current) +OUTPUT_T = 80 # future steps (8 s at 0.1 s) +POSE_DIM = 4 # x, y, cos(yaw), sin(yaw) +AGENT_STATE_DIM = 11 # x, y, cos, sin, vx, vy, width, length, is_vehicle, is_pedestrian, is_bicycle +EGO_CURRENT_STATE_DIM = 10 +STATIC_OBJECT_DIM = 10 +EGO_SHAPE_DIM = 3 # wheel_base, length, width +TURN_INDICATOR_OUTPUT_DIM = 5 +# logit order (dimensions.hpp:74-79); the published command is the index for 0..3 (TurnIndicatorsCommand) +TURN_INDICATOR_LABELS = ("NONE", "DISABLE", "ENABLE_LEFT", "ENABLE_RIGHT", "KEEP") +TURN_INDICATOR_OUTPUT_KEEP = 4 +TURN_INDICATOR_COMMAND_NAMES = {0: "NO_COMMAND", 1: "DISABLE", 2: "ENABLE_LEFT", 3: "ENABLE_RIGHT"} +TURN_INDICATORS_REPORT_DISABLE = 1 # TurnIndicatorsReport::DISABLE, the prev_report default (core.cpp:635-637) + +# ---- the encoder's 564 scene tokens, in concatenation order (T4M/model/module/encoder.py:251-265; SPEC 2.5) ------- +TOKEN_LAYOUT: Tuple[Tuple[str, int], ...] = ( + ("ego", 1), + ("neighbor", MAX_NUM_NEIGHBORS), + ("static", NUM_STATIC_OBJECTS), + ("lane", NUM_SEGMENTS_IN_LANE), + ("route", NUM_SEGMENTS_IN_ROUTE), + ("polygon", NUM_POLYGONS), + ("line_string", NUM_LINE_STRINGS), + ("goal", 1), + ("ego_shape", 1), + ("turn", 1), +) +ENCODING_TOKEN_NUM = sum(n for _, n in TOKEN_LAYOUT) # 564 (dimensions.hpp:33-35) + + +def token_slices() -> Dict[str, slice]: + """``{category: slice of the 564 tokens}``.""" + out, start = {}, 0 + for name, n in TOKEN_LAYOUT: + out[name] = slice(start, start + n) + start += n + return out + + +TOKEN_SLICES = token_slices() + +# class ids of the 14-dim positional feature (T4M encoder.py:10-21). The turn-indicator token reuses the ego-shape +# id 8 (FloatsEncoder hard-codes CLASS_TYPE_EGO_SHAPE, encoder.py:808; confirmed by the ONNX ConstantOfShape values). +POS_CLASS = {"ego": 0, "neighbor": 1, "static": 2, "lane": 3, "route": 4, "polygon": 5, "line_string": 6, + "goal": 7, "ego_shape": 8, "turn": 8} +POS_CLASS_NUM = 10 +POS_FEATURE_DIM = 4 + POS_CLASS_NUM # 14 + +# ---- in-graph pre-processing of the encoder (SPEC 3.8; T4M encoder.py) ------------------------------------------ +EGO_HISTORY_KEEP = slice(0, 6) # the 6 OLDEST ego samples are kept, rows 6..30 zeroed (encoder.py:170-175) +NEIGHBOR_HISTORY_KEEP = slice(INPUT_T + 1 - 6, INPUT_T + 1) # the 6 newest neighbour samples (rows 25..30) +TURN_INDICATOR_HISTORY = INPUT_T # turn_indicators[:, :-1]: the 30 values before the current report (encoder.py:209) +LANE_POS_INDEX = POINTS_PER_SEGMENT // 2 # 10: point used for the lane / route position feature (encoder.py:592) +POLYGON_POS_INDEX = POINTS_PER_POLYGON // 2 # 20 (LineEncoder, encoder.py:685) +LINE_STRING_POS_INDEX = POINTS_PER_LINE_STRING // 2 # 10 +LANE_FEATURE_DIM = 8 # x, y, dx, dy, left - centre (x, y), right - centre (x, y) +LANE_ATTRIBUTE_DIM = SEGMENT_POINT_DIM - LANE_FEATURE_DIM # 25: traffic light (5) + line types (2 x 10) of point 0 +NEIGHBOR_FEATURE_DIM = 9 # x, y, cos, sin, 0, 0 (velocities zeroed), width, length, valid-step flag +POLYGON_FEATURE_DIM = 2 + POLYGON_TYPE_NUM + 2 # x, y, type, dx, dy +LINE_STRING_FEATURE_DIM = 2 + LINE_STRING_TYPE_NUM + 2 # x, y, stop_line, road_border, dx, dy + +# ---- network constants (ONNX, SPEC 4.1-4.4) ------------------------------------------------------------------------ +LN_EPS = 1e-5 # every LayerNormalization of the three graphs (109 nodes; ttnn's default is 1e-12: pass it explicitly) +NUM_HEADS = 8 +HEAD_DIM = HIDDEN_DIM // NUM_HEADS # 32 +ATTN_SCALE = np.float32(0.17677669) # ONNX constant Mul_3 / Mul_2 = 1/sqrt(32) (applied to Q before Q.K^T) +MIXER_TOKENS = 64 # token_pre_project out, tokens_mlp width +MIXER_CHANNELS = 128 # channel_pre_project out, channels_mlp width +MIXER_DEPTH = 6 # encoder_mixer_depth (HFD param.json) +FUSION_DEPTH = 6 # encoder_fusion_depth +FUSION_MLP_DIM = 4 * HIDDEN_DIM +DIT_DEPTH = 3 # decoder_depth +DIT_MLP_DIM = 4 * HIDDEN_DIM +DIT_INPUT_DIM = (OUTPUT_T + 1) * POSE_DIM # 324 = preproj input / final projection output +DIT_TIME_DIM = OUTPUT_T + 1 # 81 = t_embedder input (one diffusion time per trajectory point) +TURN_HEAD_STEPS = tuple(range(1, OUTPUT_T, 10)) # final_x0[0, 1::10, :2] -> 8 points, 16 values (SPEC 4.4) + +# ---- DPM-Solver++(2M) (PKG/src/inference/solver/dpm_solver.cpp:29-33; PKG/config/diffusion_planner.param.yaml) ----- +# yaml l.17 model.multi_step_model.dpm_solver_steps: NFE = steps + 1 (COMPILE: the loop length of the trace) +DPM_SOLVER_STEPS = 10 +DPM_SOLVER_ORDER = 2 +NOISE_SCHEDULE_T = np.float32(1.0) +NOISE_SCHEDULE_TOTAL_N = np.float32(1000.0) +NOISE_SCHEDULE_BETA0 = np.float32(0.1) +NOISE_SCHEDULE_BETA1 = np.float32(20.0) +# SPEC 5.1 (recomputed from dpm_solver.cpp:80-94): the 11 solver timesteps for steps = 10, float32 +DPM_TIMESTEPS_STEPS10 = (1.0, 0.89912426, 0.78557116, 0.65344393, 0.49344844, 0.30464348, 0.14064588, 0.05360911, + 0.01809745, 0.00499264, 0.00100062) + +# ---- node parameters this port reproduces (PKG/config/diffusion_planner.param.yaml, effective YAML defaults) ------- +VELOCITY_SMOOTHING_WINDOW = 8 # yaml l.32 +STOPPING_THRESHOLD = 0.3 # yaml l.33, m/s +TURN_INDICATOR_KEEP_OFFSET = -1.25 # yaml l.34 +TURN_INDICATOR_HOLD_DURATION_S = 1.0 # yaml l.35 (stateful; see host.postprocess.TurnIndicatorManager) +TRAJECTORY_DT = 0.1 # postprocessing_utils.cpp:370 (constexpr double dt) +DELAY_STEP_MAX = OUTPUT_T // 2 # core.cpp:437 clamps delay_step to [0, 40] + +# ---- raw input tensors: the InputDataMap of DiffusionPlannerCore::create_input_data (core.cpp:414-595), batch 1 ---- +# All float32 and in the ego (base_link) frame, BEFORE normalization. ``sampled_trajectories`` is already in the +# normalised state space (x_T: zeros with the default temperature 0), ``delay`` is only read by the single-step graph. +# The speed-limit masks (lanes_has_speed_limit, route_lanes_has_speed_limit) are not inputs: the inference backend +# derives them from the normalised speed limits (PKG/include/.../inference/utils.hpp:112-123). +INPUT_SHAPES: Dict[str, Tuple[int, ...]] = { + "sampled_trajectories": (1, MAX_NUM_AGENTS, OUTPUT_T + 1, POSE_DIM), + "ego_agent_past": (1, INPUT_T + 1, POSE_DIM), + "ego_current_state": (1, EGO_CURRENT_STATE_DIM), + "neighbor_agents_past": (1, MAX_NUM_NEIGHBORS, INPUT_T + 1, AGENT_STATE_DIM), + "static_objects": (1, NUM_STATIC_OBJECTS, STATIC_OBJECT_DIM), + "lanes": (1, NUM_SEGMENTS_IN_LANE, POINTS_PER_SEGMENT, SEGMENT_POINT_DIM), + "lanes_speed_limit": (1, NUM_SEGMENTS_IN_LANE, 1), + "route_lanes": (1, NUM_SEGMENTS_IN_ROUTE, POINTS_PER_SEGMENT, SEGMENT_POINT_DIM), + "route_lanes_speed_limit": (1, NUM_SEGMENTS_IN_ROUTE, 1), + "polygons": (1, NUM_POLYGONS, POINTS_PER_POLYGON, 2 + POLYGON_TYPE_NUM), + "line_strings": (1, NUM_LINE_STRINGS, POINTS_PER_LINE_STRING, 2 + LINE_STRING_TYPE_NUM), + "goal_pose": (1, POSE_DIM), + "ego_shape": (1, EGO_SHAPE_DIM), + "turn_indicators": (1, INPUT_T + 1), + "delay": (1, 1), +} +INPUT_NAMES = tuple(INPUT_SHAPES) +# name -> (shape, dtype): the ModelBase.INPUT_SCHEMA of the API and the server (ttaw API.md section 10) +INPUT_SCHEMA = {name: (shape, np.float32) for name, shape in INPUT_SHAPES.items()} +# preprocessing_utils.cpp:34-84 normalises every tensor except these (core.cpp / node.cpp:633-634) +SKIP_NORMALIZATION = ("ego_shape", "sampled_trajectories", "turn_indicators", "delay") +# the inputs of the encoder graph, in its input order (HFD diffusion_planner_encoder.onnx) +ENCODER_INPUTS = ("ego_agent_past", "neighbor_agents_past", "static_objects", "lanes", "lanes_speed_limit", + "lanes_has_speed_limit", "route_lanes", "route_lanes_speed_limit", "route_lanes_has_speed_limit", + "polygons", "line_strings", "goal_pose", "ego_shape", "turn_indicators") + +# ---- weight files (HFD; BUNDLE_CONVENTIONS.md section 12) ----------------------------------------------------------- +ENCODER_ONNX = "diffusion_planner_encoder.onnx" +DECODER_ONNX = "diffusion_planner_decoder.onnx" +TURN_INDICATOR_ONNX = "diffusion_planner_turn_indicator.onnx" +PARAM_JSON = "diffusion_planner.param.json" +WEIGHT_MAJOR_VERSION = 5 # PKG/include/autoware/diffusion_planner/constants.hpp:22 (arg_reader.hpp:54-78) +FILE_SHA256 = { # SPEC 1.4 + ENCODER_ONNX: "2856886a3ed63b963cb18b457876d0bbd8916d345549ee5f199ceb49e3fcca49", + DECODER_ONNX: "eb30c0c0c8e80b8460d293d600c7b8c9e21f615ff6ce8de5ed01670c3d1057ca", + TURN_INDICATOR_ONNX: "07acfb58a5de1f25587fa86deb39a6de5605f8d93518e97f960ee27145f4a732", + PARAM_JSON: "ee3145b68fd1e1e44e532933dfe66cfee4384fbd637382c87ab5190c66a8e268", +} diff --git a/code/tt_diffusion_planner/reference/goldens.py b/code/tt_diffusion_planner/reference/goldens.py new file mode 100644 index 0000000000000000000000000000000000000000..38a029028ad20cb589ae0d8481682eb36052e6b9 --- /dev/null +++ b/code/tt_diffusion_planner/reference/goldens.py @@ -0,0 +1,162 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Golden tensors of the CPU reference: per-module taps and final outputs of one planning scene. + +Two products per scene (``code/scripts/ref_golden.py`` writes both): + +- the **full goldens** (``/.npz``, 10-25 MB, kept OUT of the bundle, by default under + ``research/diffusion-planner/goldens``): the raw inputs, the host features, every encoder tap (mixer taps compacted + to the valid entities, with their row indices), the encoding, for each of the 11 decoder evaluations its input + ``x`` (all 321 agents, for teacher forcing), its time, its block taps and output (valid agents), the solver + iterates, ``final_x0``, the turn pool and logits, the post-processed outputs, and the port's rewrite constants + (adaLN tables, pad-relative island constants); +- the **small goldens** (``tests/goldens/.outputs.npz``, ~0.1 MB, shipped): ``final_x0`` of the valid agents, + the logits, the ego trajectory columns, the turn command and the ego row of every solver iterate. + +Tap names follow ``reference.model.TAP_NAMES``; arrays of valid entities carry ``.rows`` (indices into the full +tensor). Device tests compare TT replay outputs against these, never TT against TT. +""" +from __future__ import annotations + +import datetime as _dt +from pathlib import Path +from typing import Any, Dict, Optional + +import numpy as np + +from . import config as C +from ..host import pipeline as hp +from ..host.postprocess import denoising_steps_ego +from ..host.solver import solver_plan +from ..ttaw.golden import TapRegistry, save_goldens +from . import rewrites as R +from .pipeline import ReferencePlanner + +__all__ = ["scene_goldens", "small_goldens", "lite_goldens", "write_scene", "SMALL_KEYS"] + +SMALL_KEYS = ("final_x0", "final_x0.rows", "logit", "trajectory", "turn_command", "denoising_ego") +MIXER_CATS = ("ego", "neighbor", "lane", "route", "polygon", "line_string") + + +def _rows(valid: np.ndarray) -> np.ndarray: + return np.flatnonzero(np.asarray(valid, bool)).astype(np.int32) + + +def scene_goldens(ref: ReferencePlanner, raw: Any, *, params: Optional[Dict[str, Any]] = None) -> Dict[str, np.ndarray]: + """Run the reference on one scene and return the full golden dict (numpy).""" + taps = TapRegistry() + res = ref.run(raw, taps=taps, keep_eval_io=True) + prep = res.prepared + f = prep.features + g: Dict[str, np.ndarray] = {} + for k, v in prep.raw.items(): + g[f"in.{k}"] = v + # host features (what the device consumes) + for name in ("ego", "neighbor", "neighbor_type", "static", "lane", "lane_attr", "lane_speed", "lane_has_speed", + "route", "route_attr", "route_speed", "route_has_speed", "polygon", "line_string", "goal", + "ego_shape", "turn", "token_valid", "key_valid", "pos"): + g[f"host.{name}"] = np.asarray(getattr(f, name)) + for cat, v in f.valid.items(): + g[f"host.valid.{cat}"] = np.asarray(v, bool) + g["host.agent_valid"] = prep.decoder.agent_valid + g["host.current_states"] = prep.decoder.current_states + t = taps.to_dict() + # encoder taps: mixer internals compacted to the valid entities + for cat in MIXER_CATS: + rows = _rows(f.valid[cat]) + for suffix in ("pre", "mixer"): + g[f"enc.{cat}.{suffix}"] = t[f"enc.{cat}.{suffix}"][rows] + g[f"enc.{cat}.{suffix}.rows"] = rows + for name, _ in C.TOKEN_LAYOUT: + g[f"enc.{name}"] = t[f"enc.{name}"] + g["enc.tokens"] = t["enc.tokens"] + for i in range(C.FUSION_DEPTH): + g[f"enc.fusion.{i}"] = t[f"enc.fusion.{i}"] + g["enc.encoding"] = res.encoding + # decoder evaluations: inputs for teacher forcing, outputs and block taps on the valid agents + arows = _rows(prep.decoder.agent_valid) + g["dec.rows"] = arows + g["dec.t"] = np.asarray(res.eval_times, np.float32) + g["dec.x_in"] = np.stack(res.eval_inputs).astype(np.float32) # [11, 321, 81, 4] + g["dec.out"] = np.stack(res.eval_outputs)[:, arows].astype(np.float32) + for k in range(len(res.eval_times)): + g[f"dec.{k}.temb"] = t[f"dec.{k}.temb"][0] + g[f"dec.{k}.x"] = t[f"dec.{k}.x"][arows] + for i in range(C.DIT_DEPTH): + g[f"dec.{k}.block{i}"] = t[f"dec.{k}.block{i}"][arows] + g["solver.iterates"] = np.stack(res.denoising_steps).astype(np.float32)[:, arows] + g["solver.timesteps"] = np.asarray(res.denoising_timesteps, np.float32) + g["final_x0"] = res.final_x0 + g["turn.pool"] = t["turn.pool"] + g["turn.logit"] = res.logit + # post-processed outputs (the API's Trajectory) + p = {k: spec[3] for k, spec in hp.RUNTIME_PARAMS.items()} + p.update(params or {}) + out = hp.make_output(res.final_x0, res.logit, prep, ref.normalization, p, denoising_steps=res.denoising_steps) + g["out.trajectory"] = out.poses + g["out.predicted_agents"] = out.predicted_agents + g["out.turn_command"] = np.asarray(out.turn_indicator["command"], np.int32) + g["out.denoising_ego"] = denoising_steps_ego(np.stack(res.denoising_steps), *ref.normalization.state()) + # the port's rewrite constants (built in float64, rounded once) + plan = solver_plan(C.DPM_SOLVER_STEPS) + tab = R.adaln_tables(ref.weights.params, plan.eval_times) + g["port.adaln.temb"] = tab.temb + for i, blk in enumerate(tab.blocks): + for k, v in blk.items(): + g[f"port.adaln.block{i}.{k}"] = v + g["port.adaln.final_gamma"], g["port.adaln.final_beta"] = tab.final_gamma, tab.final_beta + for cat in ("ego", "neighbor"): + cst = R.island_constants(ref.weights.params, cat) + for k in ("gelu_b1", "c0", "t1_pad", "g_pad", "t2_pad", "w_t1_valid"): + g[f"port.island.{cat}.{k}"] = getattr(cst, k) + g["port.solver"] = np.asarray([[u.order, u.t, u.a, u.b, u.c, u.r0] for u in plan.updates], np.float64) + return g + + +LITE_PREFIXES = ("host.valid.", "host.token_valid", "host.agent_valid", "enc.", "dec.rows", "dec.t", "final_x0", + "turn.", "out.", "solver.timesteps") + + +def lite_goldens(g: Dict[str, np.ndarray]) -> Dict[str, np.ndarray]: + """The compact per-scene goldens of the public-data instants (~1 MB): validity, the encoder category outputs and + the encoding on valid rows, final x0 of the valid agents, logits and the post-processed outputs (no inputs: they + stay in ``public_data/inputs``; no mixer or per-evaluation decoder taps).""" + out: Dict[str, np.ndarray] = {} + tok = np.flatnonzero(g["host.token_valid"]) + for k, v in g.items(): + if not k.startswith(LITE_PREFIXES) or k.startswith(("enc.fusion.", "enc.tokens")): + continue + if k.endswith((".pre", ".mixer", ".pre.rows", ".mixer.rows")): + continue + out[k] = v + for name, _ in C.TOKEN_LAYOUT: + rows = np.flatnonzero(g[f"host.valid.{name}"]) + out[f"enc.{name}"], out[f"enc.{name}.rows"] = g[f"enc.{name}"][rows], rows.astype(np.int32) + out["enc.encoding"], out["enc.encoding.rows"] = g["enc.encoding"][tok], tok.astype(np.int32) + out["final_x0"] = g["final_x0"][g["dec.rows"]] + return out + + +def small_goldens(g: Dict[str, np.ndarray]) -> Dict[str, np.ndarray]: + """The shipped subset of a full golden dict (see the module docstring).""" + rows = g["dec.rows"] + return {"final_x0": g["final_x0"][rows], "final_x0.rows": rows, "logit": g["turn.logit"], + "trajectory": g["out.trajectory"], "turn_command": g["out.turn_command"], + "denoising_ego": g["out.denoising_ego"]} + + +def write_scene(ref: ReferencePlanner, raw: Any, scene: str, full_dir: Optional[Path], small_dir: Optional[Path], + meta: Optional[Dict[str, Any]] = None, *, lite: bool = False) -> Dict[str, Any]: + g = scene_goldens(ref, raw) + info = {"scene": scene, "created": _dt.datetime.now(_dt.timezone.utc).isoformat(timespec="seconds"), + "weights_sha256": ref.weights.sha256, "solver_steps": C.DPM_SOLVER_STEPS, + "valid_counts": {k: int(np.asarray(v).sum()) for k, v in g.items() if k.startswith("host.valid.")}, + "producer": "tt_diffusion_planner.reference (fp32 CPU, torch)", **(meta or {})} + paths = {} + if full_dir is not None: + tensors = lite_goldens(g) if lite else g + info["lite"] = bool(lite) + paths["full"] = str(save_goldens(Path(full_dir) / f"{scene}.npz", tensors, info, compress=True)) + if small_dir is not None: + paths["small"] = str(save_goldens(Path(small_dir) / f"{scene}.outputs.npz", small_goldens(g), info, + compress=True)) + return {"info": info, "paths": paths, "goldens": g} diff --git a/code/tt_diffusion_planner/reference/model.py b/code/tt_diffusion_planner/reference/model.py new file mode 100644 index 0000000000000000000000000000000000000000..67cffbd4ff5883c3e3d97dc0e5a968525079403f --- /dev/null +++ b/code/tt_diffusion_planner/reference/model.py @@ -0,0 +1,226 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Pure-PyTorch fp32 re-implementation of the deployed v5.0 graphs (encoder, DiT decoder, turn-indicator head). + +Semantics follow the exported ONNX (SPEC 4.2-4.4; tier4/Diffusion-Planner ``encoder.py`` / ``mixer.py`` / ``dit.py`` +/ ``decoder.py`` read for reference, never imported), operation by operation where the order matters for float32: +attention scales Q before ``Q.K^T`` and adds the ``-inf`` key bias before the softmax; the fusion takes Q from +``LN1(x)`` but K and V from the un-normalised ``x``; SiLU is ``x * sigmoid(x)``; GELU is exact (erf) everywhere in +the encoder and in the decoder's ``preproj`` / ``t_embedder``, tanh-approximate in the six DiT MLPs and the final +projection; every LayerNorm uses epsilon 1e-5. + +Inputs are the host features of :mod:`tt_diffusion_planner.host.features` (computed from normalised inputs), so +this module holds only the network. Weights are the canonical dict of :mod:`.weights` (``w`` is ``[in, out]``). +Every module output can be recorded with a ``ttaw.golden.TapRegistry`` (names in :data:`TAP_NAMES`); valid-row +compaction of the taps is left to the caller. +""" +from __future__ import annotations + +from typing import Dict, Mapping, Optional, Tuple + +import numpy as np +import torch +import torch.nn.functional as F + +from . import config as C +from ..ttaw.golden import NULL_TAPS, TapRegistry + +__all__ = ["Encoder", "Decoder", "TurnHead", "torch_params", "TAP_NAMES"] + +Tensor = torch.Tensor + +TAP_NAMES = { + "enc..pre": "mixer categories: token_pre_project output transposed back to [E, 64, 128]", + "enc..mixer": "mixer categories: output of the 6 MixerBlocks [E, 64, 128]", + "enc.": "category output [E, 256] (masked; route with its position embedding), the fusion input rows", + "enc.tokens": "fusion input [564, 256] (category outputs + masked positional embedding)", + "enc.fusion.": "fusion block i output [564, 256]", + "enc.encoding": "final LayerNorm [564, 256] (the ``encoding`` graph output)", + "dec..temb": "t_embedder output [321, 256] of evaluation k", + "dec..x": "preproj + agent embedding [321, 256]", + "dec..block": "DiT block i output [321, 256]", + "dec..out": "model_output [321, 81, 4]", + "turn.pool": "mean of the encoding over the 564 tokens [256]", + "turn.logit": "turn_indicator_logit [5]", +} + + +def torch_params(params: Mapping[str, np.ndarray], dtype: torch.dtype = torch.float32) -> Dict[str, Tensor]: + """``{name: tensor}`` copies (the ONNX-backed arrays are read-only).""" + return {k: torch.from_numpy(np.array(v, copy=True)).to(dtype) for k, v in params.items()} + + +def _t(a, dtype=torch.float32) -> Tensor: + if isinstance(a, torch.Tensor): + return a.to(dtype) + return torch.from_numpy(np.ascontiguousarray(a)).to(dtype) + + +class _Ops: + """Shared building blocks over the canonical parameter dict.""" + + def __init__(self, p: Mapping[str, Tensor]): + self.p = p + + def linear(self, x: Tensor, name: str) -> Tensor: + return torch.matmul(x, self.p[f"{name}.w"]) + self.p[f"{name}.b"] + + def ln(self, x: Tensor, name: str) -> Tensor: + return F.layer_norm(x, x.shape[-1:], self.p[f"{name}.gamma"], self.p[f"{name}.beta"], eps=C.LN_EPS) + + def mlp(self, x: Tensor, name: str, approximate: str = "none") -> Tensor: + return self.linear(F.gelu(self.linear(x, f"{name}.fc1"), approximate=approximate), f"{name}.fc2") + + def attention(self, q_in: Tensor, kv: Tensor, name_q: str, key_bias: Optional[Tensor]) -> Tensor: + """Multi-head attention with 8 heads of 32: ``q_in`` [Sq, 256] -> Q (``name_q`` linear); ``kv`` [Sk, 512] is + the already projected K|V; ``key_bias`` [Sk] is 0 / -inf. Returns the heads concatenated [Sq, 256] + (before the output projection).""" + q = self.linear(q_in, name_q) + return self.attend(q, kv[:, :C.HIDDEN_DIM], kv[:, C.HIDDEN_DIM:], key_bias) + + @staticmethod + def attend(q: Tensor, k: Tensor, v: Tensor, key_bias: Optional[Tensor]) -> Tensor: + sq, sk = q.shape[0], k.shape[0] + qh = q.reshape(sq, C.NUM_HEADS, C.HEAD_DIM).transpose(0, 1) * float(C.ATTN_SCALE) # [H, Sq, 32] + kh = k.reshape(sk, C.NUM_HEADS, C.HEAD_DIM).permute(1, 2, 0) # [H, 32, Sk] + vh = v.reshape(sk, C.NUM_HEADS, C.HEAD_DIM).transpose(0, 1) # [H, Sk, 32] + scores = torch.matmul(qh, kh) + if key_bias is not None: + scores = scores + key_bias.reshape(1, 1, sk) + o = torch.matmul(torch.softmax(scores, dim=-1), vh) # [H, Sq, 32] + return o.transpose(0, 1).reshape(sq, C.HIDDEN_DIM) + + +def _key_bias(valid: np.ndarray) -> Tensor: + b = np.where(np.asarray(valid, bool), np.float32(0.0), np.float32(-np.inf)).astype(np.float32) + return torch.from_numpy(b) + + +class Encoder(_Ops): + """``diffusion_planner_encoder.onnx`` on host features -> ``encoding`` [564, 256].""" + + def mixer_trunk(self, x: Tensor, mod: str, taps: TapRegistry, cat: str) -> Tensor: + """channel_pre (C_in -> 128 -> 128), token_pre over the time / point axis (T -> 64 -> 64), 6 MixerBlocks, + mean over the 64 tokens -> [E, 128].""" + N = f"encoder.{mod}" + x = self.mlp(x, f"{N}.channel_pre_project") # [E, T, 128] + x = self.mlp(x.transpose(1, 2), f"{N}.token_pre_project").transpose(1, 2) # [E, 64, 128] + taps.tap(f"enc.{cat}.pre", x) + for i in range(C.MIXER_DEPTH): + B = f"{N}.blocks.{i}" + y = self.mlp(self.ln(x, f"{B}.norm1").transpose(1, 2), f"{B}.tokens_mlp").transpose(1, 2) + x = x + y + x = x + self.mlp(self.ln(x, f"{B}.norm2"), f"{B}.channels_mlp") + taps.tap(f"enc.{cat}.mixer", x) + return x.mean(dim=1) + + def head(self, x: Tensor, mod: str) -> Tensor: + """LayerNorm(128) + emb_project (128 -> 256 -> 256).""" + return self.mlp(self.ln(x, f"encoder.{mod}.norm"), f"encoder.{mod}.emb_project") + + def lanes(self, x: Tensor, attr: Tensor, speed: Tensor, has_speed: Tensor, valid: Tensor, mod: str, + taps: TapRegistry, cat: str) -> Tensor: + N = f"encoder.{mod}" + h = self.mixer_trunk(x, mod, taps, cat) + speed_emb = torch.where(has_speed, self.linear(speed, f"{N}.speed_limit_emb"), + self.p[f"{N}.unknown_speed_emb"].reshape(1, -1)) + h = h + speed_emb + self.linear(attr, f"{N}.attribute_emb") + return self.head(h, mod) * valid + + def small(self, x: Tensor, mod: str) -> Tensor: + """goal / ego-shape / turn-indicator encoders: channel MLP -> LayerNorm -> projection.""" + return self.head(self.mlp(x, f"encoder.{mod}.channel_pre_project"), mod) + + def forward(self, f, taps: TapRegistry = NULL_TAPS) -> Tensor: + valid = {k: _t(v.astype(np.float32)).reshape(-1, 1) for k, v in f.valid.items()} + out: Dict[str, Tensor] = {} + e = self.mixer_trunk(_t(f.ego)[None], "ego_encoder", taps, "ego") + out["ego"] = self.head(e, "ego_encoder") + n = self.mixer_trunk(_t(f.neighbor), "neighbor_encoder", taps, "neighbor") + n = n + self.linear(_t(f.neighbor_type), "encoder.neighbor_encoder.type_emb") + out["neighbor"] = self.head(n, "neighbor_encoder") * valid["neighbor"] + out["static"] = self.mlp(_t(f.static), "encoder.static_encoder.projection") * valid["static"] + out["lane"] = self.lanes(_t(f.lane), _t(f.lane_attr), _t(f.lane_speed), torch.from_numpy(f.lane_has_speed), + valid["lane"], "lane_encoder", taps, "lane") + route = self.lanes(_t(f.route), _t(f.route_attr), _t(f.route_speed), torch.from_numpy(f.route_has_speed), + valid["route"], "route_encoder", taps, "route") + out["route"] = route + self.p["encoder.route_position_embedding"] * valid["route"] + poly = self.mixer_trunk(_t(f.polygon), "polygon_encoder", taps, "polygon") + out["polygon"] = self.head(poly, "polygon_encoder") * valid["polygon"] + ls = self.mixer_trunk(_t(f.line_string), "line_string_encoder", taps, "line_string") + out["line_string"] = self.head(ls, "line_string_encoder") * valid["line_string"] + out["goal"] = self.small(_t(f.goal)[None], "goal_pose_encoder") + out["ego_shape"] = self.small(_t(f.ego_shape)[None], "ego_shape_encoder") + out["turn"] = self.small(_t(f.turn)[None], "turn_indicator_encoder") + for name, _ in C.TOKEN_LAYOUT: + taps.tap(f"enc.{name}", out[name]) + x = torch.cat([out[name] for name, _ in C.TOKEN_LAYOUT], dim=0) # [564, 256] + pos = self.linear(_t(f.pos), "encoder.pos_emb") * _t(f.token_valid.astype(np.float32)).reshape(-1, 1) + x = taps.tap("enc.tokens", x + pos) + key_bias = _key_bias(f.key_valid) + for i in range(C.FUSION_DEPTH): + B = f"encoder.fusion.blocks.{i}" + kv = self.linear(x, f"{B}.attn.kv") # K and V from the un-normalised x + heads = self.attention(self.ln(x, f"{B}.norm1"), kv, f"{B}.attn.q", key_bias) + x = x + self.linear(heads, f"{B}.attn.out") + x = x + self.mlp(self.ln(x, f"{B}.norm2"), f"{B}.mlp") + taps.tap(f"enc.fusion.{i}", x) + return taps.tap("enc.encoding", self.ln(x, "encoder.fusion.norm")) + + +class Decoder(_Ops): + """``diffusion_planner_decoder.onnx``: one DiT evaluation (x0 prediction) per call. + + ``cross_kv(encoding)`` computes the K|V of the three cross-attention blocks once per plan (the decoder graph + recomputes the same product at every call; hoisting it is exact, SPEC 4.6.1).""" + + def cross_kv(self, encoding: Tensor) -> Tuple[Tensor, ...]: + return tuple(self.linear(encoding, f"decoder.dit.blocks.{i}.cross_attn.kv") for i in range(C.DIT_DEPTH)) + + def time_embedding(self, t, agents: int = C.MAX_NUM_AGENTS) -> Tensor: + """t_embedder of the 81 per-point diffusion times [P, 81] (uniform in the multi-step mode) -> [P, 256].""" + tt = torch.full((agents, C.DIT_TIME_DIM), float(np.float32(t))) if np.ndim(t) == 0 else _t(t) + return self.mlp(tt, "decoder.dit.t_embedder") + + def forward(self, x_t: np.ndarray, t, kv: Tuple[Tensor, ...], agent_valid: np.ndarray, + taps: TapRegistry = NULL_TAPS, prefix: str = "dec") -> Tensor: + """``x_t`` [P, 81, 4] with P = 321 as exported, or any agent bucket ``P >= 1 + valid neighbours`` (ego first; + exact for the agents present, SPEC 4.6.4); ``agent_valid`` [P]; ``t`` a scalar time (or [P, 81]).""" + P = int(np.shape(x_t)[0]) + x = self.mlp(_t(x_t).reshape(P, C.DIT_INPUT_DIM), "decoder.dit.preproj") + c = taps.tap(f"{prefix}.temb", self.time_embedding(t, P)) + emb = self.p["decoder.dit.agent_embedding"] + x = x + torch.cat([emb[0:1], emb[1:2].expand(P - 1, -1)], dim=0) + x = taps.tap(f"{prefix}.x", x) + silu_c = c * torch.sigmoid(c) + key_bias = _key_bias(agent_valid) + for i in range(C.DIT_DEPTH): + B = f"decoder.dit.blocks.{i}" + mod = self.linear(silu_c, f"{B}.adaLN_modulation") + shift_msa, scale_msa, gate_msa, shift_mlp, scale_mlp, gate_mlp = torch.split(mod, C.HIDDEN_DIM, dim=-1) + h = self.ln(x, f"{B}.norm1") * (scale_msa + 1.0) + shift_msa + qkv = self.linear(h, f"{B}.attn.qkv") + q, k, v = torch.split(qkv, C.HIDDEN_DIM, dim=-1) + x = x + gate_msa * self.linear(self.attend(q, k, v, key_bias), f"{B}.attn.out") + h = self.ln(x, f"{B}.norm2") * (scale_mlp + 1.0) + shift_mlp + x = x + gate_mlp * self.mlp(h, f"{B}.mlp1", "tanh") + heads = self.attention(self.ln(x, f"{B}.norm3"), kv[i], f"{B}.cross_attn.q", None) # no key mask + x = x + self.linear(heads, f"{B}.cross_attn.out") + x = x + self.mlp(self.ln(x, f"{B}.norm4"), f"{B}.mlp2", "tanh") + taps.tap(f"{prefix}.block{i}", x) + Fn = "decoder.dit.final_layer" + shift, scale = torch.split(self.linear(silu_c, f"{Fn}.adaLN_modulation"), C.HIDDEN_DIM, dim=-1) + h = self.ln(x, f"{Fn}.norm_final") * (scale + 1.0) + shift + h = F.gelu(self.linear(self.ln(h, f"{Fn}.proj.0"), f"{Fn}.proj.1"), approximate="tanh") + out = self.linear(self.ln(h, f"{Fn}.proj.3"), f"{Fn}.proj.4").reshape(P, C.OUTPUT_T + 1, C.POSE_DIM) + return taps.tap(f"{prefix}.out", out) + + +class TurnHead(_Ops): + """``diffusion_planner_turn_indicator.onnx``: ``W [272 -> 5]`` over ``final_x0[0, 1::10, :2]`` (16 values) and the + token mean of the encoding (256).""" + + def forward(self, encoding: Tensor, final_x0: np.ndarray, taps: TapRegistry = NULL_TAPS) -> Tensor: + pool = taps.tap("turn.pool", encoding.mean(dim=0)) + ego = _t(final_x0)[0, 1::10, :2].reshape(-1) + feat = torch.cat([ego, pool], dim=0)[None] + return taps.tap("turn.logit", self.linear(feat, "decoder.turn_indicator_predictor")[0]) diff --git a/code/tt_diffusion_planner/reference/ort.py b/code/tt_diffusion_planner/reference/ort.py new file mode 100644 index 0000000000000000000000000000000000000000..2c9a145cef3e6c424e32db3d3f6cfc5af5007e64 --- /dev/null +++ b/code/tt_diffusion_planner/reference/ort.py @@ -0,0 +1,137 @@ +# SPDX-License-Identifier: Apache-2.0 +"""ONNX Runtime on the shipped v5.0 ONNX files: the oracle of the CPU reference (research and tests only; +onnxruntime is not a runtime dependency of the bundle or the image). + +``OrtPlanner.run(raw)`` is the Autoware ``multi_step`` pipeline with ORT as the backend (a port of the verified +research script ``research/diffusion-planner/scripts/dp_reference.py``): the node's normalization and speed masks, +encoder.onnx, the DPM-Solver++(2M) loop of :mod:`tt_diffusion_planner.host.solver` calling decoder.onnx 11 times, +then turn_indicator.onnx. ``taps=[...]`` exposes intermediate tensors of the encoder / decoder graphs by ONNX tensor +name (the decoder taps are recorded per evaluation), for the per-module PCC test of the reference. +""" +from __future__ import annotations + +from dataclasses import dataclass, field +from pathlib import Path +from typing import Any, Dict, List, Sequence + +import numpy as np + +from . import config as C +from ..host.normalize import normalize_inputs, speed_masks +from ..host.solver import apply_prefix_constraint, dpm_solver_sample +from ..host.features import decoder_masks +from ..ttaw import io as tio +from .weights import load_param_json + +__all__ = ["OrtPlanner", "OrtResult", "available"] + + +def available() -> bool: + import importlib.util + + return importlib.util.find_spec("onnxruntime") is not None + + +def _session(model: Any, threads: int, optimize: bool = True): + import onnxruntime as ort + + so = ort.SessionOptions() + so.graph_optimization_level = (ort.GraphOptimizationLevel.ORT_ENABLE_ALL if optimize + else ort.GraphOptimizationLevel.ORT_DISABLE_ALL) + so.intra_op_num_threads = threads + so.inter_op_num_threads = 1 + so.log_severity_level = 3 + src = model.SerializeToString() if hasattr(model, "SerializeToString") else str(model) + return ort.InferenceSession(src, so, providers=["CPUExecutionProvider"]) + + +def _with_outputs(path: Path, names: Sequence[str]): + """The model with extra graph outputs (intermediate tensors by name).""" + import onnx + + m = onnx.load(str(path)) + have = {o.name for o in m.graph.output} + m.graph.output.extend([onnx.ValueInfoProto(name=n) for n in names if n not in have]) + return m + + +@dataclass +class OrtResult: + norm: Dict[str, np.ndarray] + encoding: np.ndarray # [1, 564, 256] + final_x0: np.ndarray # [1, 321, 81, 4] + logit: np.ndarray # [1, 5] + denoising_steps: List[np.ndarray] + denoising_timesteps: List[float] + eval_times: List[float] + eval_inputs: List[np.ndarray] = field(default_factory=list) + eval_outputs: List[np.ndarray] = field(default_factory=list) + taps: Dict[str, np.ndarray] = field(default_factory=dict) + + +class OrtPlanner: + """``OrtPlanner(weights_dir, threads=4, encoder_taps=(...), decoder_taps=(...))``. With taps the sessions run + without graph optimizations (every intermediate kept as computed); without, ``ORT_ENABLE_ALL`` like + ``onnxruntime_inference.cpp:172``.""" + + def __init__(self, weights_dir: Path, *, threads: int = 4, encoder_taps: Sequence[str] = (), + decoder_taps: Sequence[str] = (), turn_taps: Sequence[str] = ()): + wd = Path(weights_dir) + self.normalization = load_param_json(wd / C.PARAM_JSON) + self.encoder_taps, self.decoder_taps, self.turn_taps = list(encoder_taps), list(decoder_taps), list(turn_taps) + opt = not (encoder_taps or decoder_taps or turn_taps) + self.enc = _session(_with_outputs(wd / C.ENCODER_ONNX, self.encoder_taps), threads, opt) + self.dec = _session(_with_outputs(wd / C.DECODER_ONNX, self.decoder_taps), threads, opt) + self.turn = _session(_with_outputs(wd / C.TURN_INDICATOR_ONNX, self.turn_taps), threads, opt) + + def normalize(self, raw: Any) -> Dict[str, np.ndarray]: + arrays = tio.load_named_arrays(raw, C.INPUT_SCHEMA) + norm = normalize_inputs(arrays, self.normalization.observation) + norm.update(speed_masks(norm)) + return norm + + def encode(self, norm: Dict[str, np.ndarray]) -> Dict[str, np.ndarray]: + feed = {k: norm[k] for k in C.ENCODER_INPUTS} + outs = self.enc.run(["encoding"] + self.encoder_taps, feed) + return dict(zip(["encoding"] + self.encoder_taps, outs)) + + def decode(self, encoding: np.ndarray, x: np.ndarray, t, neighbor_agents_past: np.ndarray) -> Dict[str, Any]: + dt = np.full((1, C.MAX_NUM_AGENTS, C.OUTPUT_T + 1, 1), np.float32(t), np.float32) + feed = {"encoding": encoding, "sampled_trajectories": np.asarray(x, np.float32).reshape( + 1, C.MAX_NUM_AGENTS, C.OUTPUT_T + 1, C.POSE_DIM), "diffusion_time": dt, + "neighbor_agents_past": neighbor_agents_past} + outs = self.dec.run(["model_output"] + self.decoder_taps, feed) + return dict(zip(["model_output"] + self.decoder_taps, outs)) + + def turn_logit(self, encoding: np.ndarray, final_x0: np.ndarray) -> Dict[str, np.ndarray]: + outs = self.turn.run(["turn_indicator_logit"] + self.turn_taps, + {"encoding": encoding, "final_x0": np.asarray(final_x0, np.float32)}) + return dict(zip(["turn_indicator_logit"] + self.turn_taps, outs)) + + def run(self, raw: Any, *, steps: int = C.DPM_SOLVER_STEPS, keep_eval_io: bool = False) -> OrtResult: + norm = self.normalize(raw) + enc = self.encode(norm) + encoding = enc["encoding"] + taps: Dict[str, np.ndarray] = {f"enc:{k}": v for k, v in enc.items() if k != "encoding"} + cs = decoder_masks(norm).current_states + nb = norm["neighbor_agents_past"] + eval_in: List[np.ndarray] = [] + eval_out: List[np.ndarray] = [] + + def model_fn(x: np.ndarray, t) -> np.ndarray: + k = len(eval_out) + out = self.decode(encoding, x, t, nb) + for name in self.decoder_taps: + taps[f"dec{k}:{name}"] = out[name] + y = out["model_output"][0] + eval_out.append(y if keep_eval_io else np.empty(0)) + if keep_eval_io: + eval_in.append(np.array(x, np.float32, copy=True)) + return y + + res = dpm_solver_sample(norm["sampled_trajectories"][0], model_fn, + lambda x: apply_prefix_constraint(x, cs), steps) + tl = self.turn_logit(encoding, res.final_x[None]) + taps.update({f"turn:{k}": v for k, v in tl.items() if k != "turn_indicator_logit"}) + return OrtResult(norm, encoding, res.final_x[None], tl["turn_indicator_logit"], res.denoising_steps, + res.denoising_timesteps, res.eval_times, eval_in, eval_out if keep_eval_io else [], taps) diff --git a/code/tt_diffusion_planner/reference/pipeline.py b/code/tt_diffusion_planner/reference/pipeline.py new file mode 100644 index 0000000000000000000000000000000000000000..4049fccdb12bc81403422be601cdff8df0e34599 --- /dev/null +++ b/code/tt_diffusion_planner/reference/pipeline.py @@ -0,0 +1,123 @@ +# SPDX-License-Identifier: Apache-2.0 +"""The fp32 CPU reference planner: host pre-processing -> encoder -> DPM-Solver++(2M) over the DiT decoder (11 +evaluations) -> turn head -> host post-processing, i.e. the Autoware ``multi_step`` mode with guidance off. + + from tt_diffusion_planner.reference import ReferencePlanner + ref = ReferencePlanner() # weights: find_weights_dir() or weights_dir=... + out = ref(inputs="code/tt_diffusion_planner/samples/kashiwanoha_dense.npz") # ttaw.outputs.Trajectory + res = ref.run(raw_inputs, taps=TapRegistry()) # every intermediate (golden generation) + +``__call__`` returns exactly what the device model returns (same ``host.prepare`` / ``host.make_output``), so the +stored ``samples/.reference.json`` is this class's ``to_dict()``. +""" +from __future__ import annotations + +import time +from dataclasses import dataclass, field +from pathlib import Path +from typing import Any, Dict, List, Optional + +import numpy as np +import torch + +from . import config as C +from ..host import pipeline as hp +from ..host.solver import apply_prefix_constraint, dpm_solver_sample +from ..ttaw import io as tio +from ..ttaw.golden import NULL_TAPS, TapRegistry +from .model import Decoder, Encoder, TurnHead, torch_params +from .weights import PlannerWeights, find_weights_dir, load_weights + +__all__ = ["ReferencePlanner", "ReferenceResult"] + +MODEL_NAME = "diffusion-planner-p150" + + +@dataclass +class ReferenceResult: + prepared: hp.Prepared + encoding: np.ndarray # [564, 256] + final_x0: np.ndarray # [321, 81, 4] normalised + logit: np.ndarray # [5] + denoising_steps: List[np.ndarray] + denoising_timesteps: List[float] + eval_times: List[float] + eval_inputs: List[np.ndarray] = field(default_factory=list) # x fed to each decoder evaluation + eval_outputs: List[np.ndarray] = field(default_factory=list) # model_output of each evaluation + timing_ms: Dict[str, float] = field(default_factory=dict) + + +class ReferencePlanner: + """CPU fp32 reference of the deployed network plus the shared host code. ``threads`` bounds torch's intra-op + threads (the workspace rule is <= 4).""" + + def __init__(self, weights_dir: Optional[str] = None, *, weights: Optional[PlannerWeights] = None, + threads: Optional[int] = 4, steps: int = C.DPM_SOLVER_STEPS): + if weights is None: + wd = find_weights_dir(weights_dir) + if wd is None: + raise FileNotFoundError("Diffusion Planner v5.0 weights not found: pass weights_dir= or set " + "DIFFUSION_PLANNER_WEIGHTS_DIR") + weights = load_weights(wd) + self.weights = weights + self.normalization = weights.normalization + if threads: + torch.set_num_threads(int(threads)) + p = torch_params(weights.params) + self.encoder, self.decoder, self.turn = Encoder(p), Decoder(p), TurnHead(p) + self.steps = steps + + # ---- network --------------------------------------------------------------------------------------------- + @torch.no_grad() + def run(self, raw: Any, *, taps: TapRegistry = NULL_TAPS, keep_eval_io: bool = False) -> ReferenceResult: + """One plan on raw ONNX-named inputs (mapping, ``.npz`` path / bytes or the JSON envelope).""" + arrays = tio.load_named_arrays(raw, C.INPUT_SCHEMA) + t0 = time.perf_counter() + prep = hp.prepare(arrays, self.normalization.observation) + t1 = time.perf_counter() + encoding = self.encoder.forward(prep.features, taps) + t2 = time.perf_counter() + kv = self.decoder.cross_kv(encoding) + cs = prep.decoder.current_states + eval_in: List[np.ndarray] = [] + eval_out: List[np.ndarray] = [] + + def model_fn(x: np.ndarray, t) -> np.ndarray: + k = len(eval_out) + y = self.decoder.forward(x, t, kv, prep.decoder.agent_valid, taps, prefix=f"dec.{k}").numpy() + if keep_eval_io: + eval_in.append(x.copy()) + eval_out.append(y if keep_eval_io else np.empty(0)) + return y + + res = dpm_solver_sample(prep.x_T, model_fn, lambda x: apply_prefix_constraint(x, cs), self.steps) + t3 = time.perf_counter() + logit = self.turn.forward(encoding, res.final_x, taps).numpy() + t4 = time.perf_counter() + for k, x in enumerate(res.denoising_steps): + taps.tap(f"solver.x{k}", x) + taps.tap("final_x0", res.final_x) + return ReferenceResult(prep, encoding.numpy(), res.final_x, logit, res.denoising_steps, + res.denoising_timesteps, res.eval_times, eval_in, eval_out if keep_eval_io else [], + {"preprocess": (t1 - t0) * 1e3, "encoder": (t2 - t1) * 1e3, "solver": (t3 - t2) * 1e3, + "turn": (t4 - t3) * 1e3}) + + # ---- the API-shaped call ---------------------------------------------------------------------------------- + def __call__(self, inputs: Any, **params: Any): + """``model(inputs=...)`` of the CPU reference: the same host pre / post-processing as the device model.""" + p = {k: spec[3] for k, spec in hp.RUNTIME_PARAMS.items()} + unknown = sorted(set(params) - set(p)) + if unknown: + raise tio.InputError(f"unknown parameter(s) {unknown}; allowed: {sorted(p)}") + p.update(params) + t0 = time.perf_counter() + res = self.run(inputs) + total = (time.perf_counter() - t0) * 1e3 + steps = res.denoising_steps if p.get("return_denoising_steps") else None + return hp.make_output(res.final_x0, res.logit, res.prepared, self.normalization, p, model=MODEL_NAME, + denoising_steps=steps, timing_ms={**res.timing_ms, "total": total}, + meta={"reference": "fp32 CPU (torch)"}) + + @staticmethod + def sample_path(name: str = "kashiwanoha_dense.npz") -> Path: + return Path(__file__).resolve().parents[1] / "samples" / name diff --git a/code/tt_diffusion_planner/reference/rewrites.py b/code/tt_diffusion_planner/reference/rewrites.py new file mode 100644 index 0000000000000000000000000000000000000000..3c50dd43b665ef1075f3496c0200bd18bf4c4501 --- /dev/null +++ b/code/tt_diffusion_planner/reference/rewrites.py @@ -0,0 +1,183 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Exact rewrites of the deployed graph that the TT port applies, as constants and CPU forwards (SPEC 4.6, PLAN 2.12). + +Each rewrite is exact in real arithmetic; its constants are computed in float64 from the float32 weights and rounded +once to float32 (the "fold in fp64, round once" rule of ``ttaw.weights``). The CPU forwards here use them and are +tested against the as-exported :mod:`.model` (``tests/test_reference_host.py``), so the device port has a CPU oracle in +its own parameterisation. + +1. **Per-step adaLN tables folded into the LayerNorm affine** (SPEC 4.6.2, 8.1). With a uniform diffusion time (the + multi-step mode) the t-embedding and every adaLN output are the same for all 321 agents and depend only on t, so + for the 11 evaluation times ``modulate(LN(x; g, b), shift, scale) = LN(x; g (1 + scale), b (1 + scale) + shift)`` + and the gates are per-step ``[256]`` rows. :func:`adaln_tables` builds them for ``SolverPlan.eval_times``. +2. **Cross-attention K/V hoisted** (SPEC 4.6.1): ``encoding @ W_kv + b_kv`` of the three blocks once per plan + (the decoder graph recomputes it at each of the 11 calls). +3. **Pad-relative fp32 pre-projection island** (SPEC 4.6.7, probe P12): after the in-graph history truncation the + ego (rows 6..30) and neighbour (rows 0..24) inputs of ``channel_pre_project`` are zero rows, which all map to + ``c0 = fc2(gelu(b1)) + b2``. With the all-zero ("pad") agent's outputs ``t1_pad = b_t1 + c0 (x) sum_t W_t1[t]``, + ``g_pad = gelu(t1_pad)``, ``t2_pad = g_pad @ W_t2 + b_t2`` the island becomes + ``t1 = t1_pad + (z_valid - c0)^T @ W_t1[valid rows]`` and ``t2 = t2_pad + (gelu(t1) - g_pad) @ W_t2``, so the + device's TF32-like fp32 matmuls only see deviations from the pad agent (P12: deviation PCC 0.995706 -> 0.999991 + on ``straight``). ``gelu_b1 = gelu(b_c1)`` also lets ``channel_pre`` itself run pad-relative: + ``z - c0 = (gelu(x @ W_c1 + b_c1) - gelu_b1) @ W_c2``. +""" +from __future__ import annotations + +from dataclasses import dataclass +from typing import Dict, Mapping, Sequence, Tuple + +import numpy as np +import torch +import torch.nn.functional as F + +from . import config as C +from .model import Decoder, _t + +__all__ = ["AdaLNTables", "adaln_tables", "decoder_forward_folded", "IslandConstants", "island_constants", + "island_forward_pad_relative", "island_forward_exported", "cross_kv"] + + +# ---- 1. adaLN tables --------------------------------------------------------------------------------------------- +@dataclass +class AdaLNTables: + """float32 tables for K evaluation times: ``temb [K, 256]``; per DiT block ``i``: ``norm1_gamma``, + ``norm1_beta``, ``gate_msa``, ``norm2_gamma``, ``norm2_beta``, ``gate_mlp`` ``[K, 256]``; the final layer's + ``final_gamma`` / ``final_beta`` ``[K, 256]``.""" + + eval_times: Tuple[float, ...] + temb: np.ndarray + blocks: Tuple[Dict[str, np.ndarray], ...] + final_gamma: np.ndarray + final_beta: np.ndarray + + +def adaln_tables(params: Mapping[str, np.ndarray], eval_times: Sequence[float]) -> AdaLNTables: + p = {k: torch.from_numpy(np.array(v, np.float64)) for k, v in params.items() if k.startswith("decoder.dit")} + + def lin(x, n): + return x @ p[f"{n}.w"] + p[f"{n}.b"] + + t = torch.tensor([[float(np.float32(v))] * C.DIT_TIME_DIM for v in eval_times], dtype=torch.float64) + c = lin(F.gelu(lin(t, "decoder.dit.t_embedder.fc1")), "decoder.dit.t_embedder.fc2") # [K, 256] + silu = c * torch.sigmoid(c) + blocks = [] + for i in range(C.DIT_DEPTH): + B = f"decoder.dit.blocks.{i}" + sh_msa, sc_msa, g_msa, sh_mlp, sc_mlp, g_mlp = torch.split(lin(silu, f"{B}.adaLN_modulation"), C.HIDDEN_DIM, -1) + g1, b1 = p[f"{B}.norm1.gamma"], p[f"{B}.norm1.beta"] + g2, b2 = p[f"{B}.norm2.gamma"], p[f"{B}.norm2.beta"] + blocks.append({k: v.to(torch.float32).numpy() for k, v in { + "norm1_gamma": g1 * (1 + sc_msa), "norm1_beta": b1 * (1 + sc_msa) + sh_msa, "gate_msa": g_msa, + "norm2_gamma": g2 * (1 + sc_mlp), "norm2_beta": b2 * (1 + sc_mlp) + sh_mlp, "gate_mlp": g_mlp}.items()}) + Fn = "decoder.dit.final_layer" + sh, sc = torch.split(lin(silu, f"{Fn}.adaLN_modulation"), C.HIDDEN_DIM, -1) + gf, bf = p[f"{Fn}.norm_final.gamma"], p[f"{Fn}.norm_final.beta"] + return AdaLNTables(tuple(float(v) for v in eval_times), c.to(torch.float32).numpy(), tuple(blocks), + (gf * (1 + sc)).to(torch.float32).numpy(), (bf * (1 + sc) + sh).to(torch.float32).numpy()) + + +def cross_kv(params: Mapping[str, np.ndarray], encoding: np.ndarray) -> Tuple[np.ndarray, ...]: + """Hoisted cross-attention K|V ``[564, 512]`` of each DiT block (float32 matmul, as the exported graph).""" + enc = np.asarray(encoding, np.float32) + return tuple((enc @ params[f"decoder.dit.blocks.{i}.cross_attn.kv.w"] + + params[f"decoder.dit.blocks.{i}.cross_attn.kv.b"]).astype(np.float32) for i in range(C.DIT_DEPTH)) + + +@torch.no_grad() +def decoder_forward_folded(dec: Decoder, tables: AdaLNTables, k: int, x_t: np.ndarray, kv: Sequence[torch.Tensor], + agent_valid: np.ndarray) -> torch.Tensor: + """One decoder evaluation with the step-``k`` tables: LayerNorms with folded affine (no SiLU / adaLN matmuls), + gates as constant rows. Same result as ``Decoder.forward`` up to float32 rounding.""" + from .model import _key_bias + + P = int(np.shape(x_t)[0]) + x = dec.mlp(_t(x_t).reshape(P, C.DIT_INPUT_DIM), "decoder.dit.preproj") + emb = dec.p["decoder.dit.agent_embedding"] + x = x + torch.cat([emb[0:1], emb[1:2].expand(P - 1, -1)], dim=0) + key_bias = _key_bias(agent_valid) + + def ln_folded(h, gamma, beta): + return F.layer_norm(h, h.shape[-1:], _t(gamma), _t(beta), eps=C.LN_EPS) + + for i in range(C.DIT_DEPTH): + B, T = f"decoder.dit.blocks.{i}", tables.blocks[i] + h = ln_folded(x, T["norm1_gamma"][k], T["norm1_beta"][k]) + q, kk, v = torch.split(dec.linear(h, f"{B}.attn.qkv"), C.HIDDEN_DIM, dim=-1) + x = x + _t(T["gate_msa"][k]) * dec.linear(dec.attend(q, kk, v, key_bias), f"{B}.attn.out") + h = ln_folded(x, T["norm2_gamma"][k], T["norm2_beta"][k]) + x = x + _t(T["gate_mlp"][k]) * dec.mlp(h, f"{B}.mlp1", "tanh") + heads = dec.attention(dec.ln(x, f"{B}.norm3"), kv[i], f"{B}.cross_attn.q", None) + x = x + dec.linear(heads, f"{B}.cross_attn.out") + x = x + dec.mlp(dec.ln(x, f"{B}.norm4"), f"{B}.mlp2", "tanh") + Fn = "decoder.dit.final_layer" + h = ln_folded(x, tables.final_gamma[k], tables.final_beta[k]) + h = F.gelu(dec.linear(dec.ln(h, f"{Fn}.proj.0"), f"{Fn}.proj.1"), approximate="tanh") + return dec.linear(dec.ln(h, f"{Fn}.proj.3"), f"{Fn}.proj.4").reshape(P, C.OUTPUT_T + 1, C.POSE_DIM) + + +# ---- 3. pad-relative island --------------------------------------------------------------------------------------- +@dataclass +class IslandConstants: + """Pad-agent constants of one pre-projection island (float32, built in float64). ``valid_rows`` are the time + rows that can be non-zero after the truncation (ego 0..5, neighbours 25..30); ``w_t1_valid`` is + ``W_t1[valid_rows]`` ``[6, 64]``.""" + + category: str + valid_rows: Tuple[int, ...] + gelu_b1: np.ndarray # [128] gelu(b_c1): channel_pre's first activation of a zero row + c0: np.ndarray # [128] channel_pre of a zero row + t1_pad: np.ndarray # [128, 64] + g_pad: np.ndarray # [128, 64] + t2_pad: np.ndarray # [128, 64] + w_t1_valid: np.ndarray # [6, 64] + + +ISLAND_ROWS = {"neighbor": tuple(range(C.NEIGHBOR_HISTORY_KEEP.start, C.NEIGHBOR_HISTORY_KEEP.stop)), + "ego": tuple(range(C.EGO_HISTORY_KEEP.start, C.EGO_HISTORY_KEEP.stop))} +ISLAND_MODULE = {"neighbor": "encoder.neighbor_encoder", "ego": "encoder.ego_encoder"} + + +def island_constants(params: Mapping[str, np.ndarray], category: str) -> IslandConstants: + N = ISLAND_MODULE[category] + w = {k: torch.from_numpy(np.array(v, np.float64)) for k, v in params.items() if k.startswith(N + ".")} + gelu_b1 = F.gelu(w[f"{N}.channel_pre_project.fc1.b"]) + c0 = gelu_b1 @ w[f"{N}.channel_pre_project.fc2.w"] + w[f"{N}.channel_pre_project.fc2.b"] + w_t1 = w[f"{N}.token_pre_project.fc1.w"] # [31, 64] + t1_pad = w[f"{N}.token_pre_project.fc1.b"][None, :] + c0[:, None] * w_t1.sum(0)[None, :] + g_pad = F.gelu(t1_pad) + t2_pad = g_pad @ w[f"{N}.token_pre_project.fc2.w"] + w[f"{N}.token_pre_project.fc2.b"] + rows = ISLAND_ROWS[category] + f32 = lambda t: t.to(torch.float32).numpy() # noqa: E731 + return IslandConstants(category, rows, f32(gelu_b1), f32(c0), f32(t1_pad), f32(g_pad), f32(t2_pad), + f32(w_t1[list(rows)])) + + +def island_forward_exported(params: Mapping[str, np.ndarray], category: str, x: np.ndarray, + dtype: torch.dtype = torch.float64) -> Tuple[torch.Tensor, torch.Tensor]: + """The island as exported: ``x`` ``[E, 31, C_in]`` -> ``t1``, ``t2`` ``[E, 128, 64]`` (``dtype`` math).""" + N = ISLAND_MODULE[category] + w = {k: torch.from_numpy(np.array(v, np.float64)).to(dtype) for k, v in params.items() if k.startswith(N + ".")} + + def lin(h, n): + return h @ w[f"{N}.{n}.w"] + w[f"{N}.{n}.b"] + + z = lin(F.gelu(lin(torch.from_numpy(np.asarray(x, np.float64)).to(dtype), "channel_pre_project.fc1")), + "channel_pre_project.fc2") + t1 = lin(z.transpose(1, 2), "token_pre_project.fc1") + return t1, lin(F.gelu(t1), "token_pre_project.fc2") + + +def island_forward_pad_relative(params: Mapping[str, np.ndarray], consts: IslandConstants, x: np.ndarray, + dtype: torch.dtype = torch.float64) -> Tuple[torch.Tensor, torch.Tensor]: + """The pad-relative island on the valid rows only (``x`` ``[E, 31, C_in]``; rows outside ``valid_rows`` must be + zero, which the truncation guarantees). Matmuls see deviations from the pad agent only.""" + N = ISLAND_MODULE[consts.category] + w = {k: torch.from_numpy(np.array(v, np.float64)).to(dtype) for k, v in params.items() if k.startswith(N + ".")} + cst = {k: torch.from_numpy(np.array(getattr(consts, k), np.float64)).to(dtype) + for k in ("gelu_b1", "c0", "t1_pad", "g_pad", "t2_pad", "w_t1_valid")} + xv = torch.from_numpy(np.asarray(x, np.float64)[:, list(consts.valid_rows)]).to(dtype) # [E, 6, C_in] + h = F.gelu(xv @ w[f"{N}.channel_pre_project.fc1.w"] + w[f"{N}.channel_pre_project.fc1.b"]) - cst["gelu_b1"] + dz = h @ w[f"{N}.channel_pre_project.fc2.w"] # z - c0, [E, 6, 128] + t1 = cst["t1_pad"] + dz.transpose(1, 2) @ cst["w_t1_valid"] # [E, 128, 64] + t2 = cst["t2_pad"] + (F.gelu(t1) - cst["g_pad"]) @ w[f"{N}.token_pre_project.fc2.w"] + return t1, t2 diff --git a/code/tt_diffusion_planner/reference/weights.py b/code/tt_diffusion_planner/reference/weights.py new file mode 100644 index 0000000000000000000000000000000000000000..b4f9b15825530d4988e58310de73df83adb2583a --- /dev/null +++ b/code/tt_diffusion_planner/reference/weights.py @@ -0,0 +1,442 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Weights of the Diffusion Planner v5.0 export, read as DATA from the three ONNX files and the param JSON. + +Every tensor is addressed through the graph node that consumes it (``ttaw.weights.OnnxWeights``), never by +initializer name: most MatMul weights are anonymous (``onnx::MatMul_4039``), several biases are deduplicated across +modules (``static_encoder/projection/fc2`` reads ``...fc1.bias``; ``route_encoder/attribute_emb`` reads +``...speed_limit_emb.bias``), the decoder cross-attention K/V of all three blocks is one unnamed ``[256, 1536]`` +MatMul followed by a Split, and the fusion / cross-attention Q-K-V biases and weights are constant-folded tensors +(SPEC 6.2). The result is one flat ``{canonical name: float32 array}`` dict that the CPU reference and the ttnn +graph both consume: + +- linear layers: ``.w`` ``[in, out]`` (``y = x @ w + b``) and ``.b`` ``[out]``; +- LayerNorms: ``.gamma`` / ``.beta``; +- attention: ``...attn.q`` / ``.kv`` (fusion: Q from LN(x), K|V from x), ``...attn.qkv`` (DiT self-attention), + ``...cross_attn.q`` / ``.kv`` (K|V of one block, a column block of the fused cross K/V MatMul), ``...out``; +- embeddings: ``decoder.dit.agent_embedding`` ``[2, 256]`` (ego, neighbour), ``encoder.route_position_embedding`` + ``[25, 256]``, ``encoder._encoder.unknown_speed_emb`` ``[128]``. + +There is no BatchNorm in this network, so nothing is folded here. The exact rewrites the TT port applies on top of +these tensors (per-step adaLN tables folded into the LayerNorm affine, hoisted cross K/V, the pad-relative fp32 +pre-projection island) are in :mod:`.rewrites`, computed from this dict in float64 with one final rounding. + +The loader also checks the invariants the reference and the port rely on (LayerNorm epsilon 1e-5 everywhere, exact +GELU in the encoder and the decoder pre-projection / t-embedder, tanh GELU in the DiT MLPs and the final projection, +attention scale 1/sqrt(32), the ``-inf`` key mask), so a different export fails loudly instead of silently. +""" +from __future__ import annotations + +import json +import math +from dataclasses import dataclass, field +from pathlib import Path +from typing import Any, Dict, List, Mapping, Optional, Tuple + +import numpy as np + +from . import config as C +from ..ttaw.weights import OnnxWeights, file_sha256 + +__all__ = ["PlannerWeights", "load_weights", "load_param_json", "Normalization", "find_weights_dir", + "MIXER_ENCODERS", "SMALL_ENCODERS"] + +# categories with an MLP-Mixer trunk -> ONNX module prefix +MIXER_ENCODERS = {"ego": "ego_encoder", "neighbor": "neighbor_encoder", "lane": "lane_encoder", + "route": "route_encoder", "polygon": "polygon_encoder", "line_string": "line_string_encoder"} +# categories encoded by a channel MLP + LayerNorm + projection (goal pose, ego shape, turn indicators) +SMALL_ENCODERS = {"goal": "goal_pose_encoder", "ego_shape": "ego_shape_encoder", "turn": "turn_indicator_encoder"} + + +# ------------------------------------------------------------------------------------------------ param JSON + +@dataclass(frozen=True) +class Normalization: + """``observation_normalizer`` (per input tensor: mean / std over the last dim) and ``state_normalizer`` + (per agent and pose dim) of ``diffusion_planner.param.json`` (PKG/include/.../utils/arg_reader.hpp:80-140).""" + + observation: Dict[str, Tuple[np.ndarray, np.ndarray]] + state_mean: np.ndarray # [321, 4] (or [4]) + state_std: np.ndarray + major_version: int + args: Dict[str, Any] = field(default_factory=dict) + + def state(self) -> Tuple[np.ndarray, np.ndarray]: + """``(mean, std)`` broadcastable to ``[321, T, 4]``.""" + return _per_agent(self.state_mean), _per_agent(self.state_std) + + +def _per_agent(v: np.ndarray) -> np.ndarray: + v = np.asarray(v, np.float32).reshape(-1) + if v.size == C.POSE_DIM: + return v.reshape(1, 1, C.POSE_DIM) + if v.size == C.MAX_NUM_AGENTS * C.POSE_DIM: + return v.reshape(C.MAX_NUM_AGENTS, 1, C.POSE_DIM) + raise ValueError(f"unsupported state normalizer size {v.size}") + + +def load_param_json(path: Path) -> Normalization: + """Parse ``diffusion_planner.param.json`` like ``arg_reader.hpp``; refuses a major version other than 5.""" + with open(path) as f: + j = json.load(f) + major = int(j.get("major_version", -1)) + if major != C.WEIGHT_MAJOR_VERSION: + raise ValueError(f"{path}: major_version {major}, this port needs {C.WEIGHT_MAJOR_VERSION} (constants.hpp:22)") + obs = {} + for key, v in j["observation_normalizer"].items(): + mean = np.asarray(v.get("mean", []), np.float32).reshape(-1) + std = np.asarray(v.get("std", []), np.float32).reshape(-1) + if mean.shape != std.shape: + raise ValueError(f"{path}: normalizer {key!r} mean / std sizes differ") + obs[key] = (mean, std) + sn = j["state_normalizer"] + args = {k: v for k, v in j.items() if k not in ("observation_normalizer", "state_normalizer")} + return Normalization(obs, np.asarray(sn["mean"], np.float32), np.asarray(sn["std"], np.float32), major, args) + + +# ------------------------------------------------------------------------------------------------ ONNX weights + +@dataclass +class PlannerWeights: + """The flat canonical parameter dict plus provenance (file sha256) and the export facts checked at load.""" + + params: Dict[str, np.ndarray] + normalization: Normalization + sha256: Dict[str, str] + facts: Dict[str, Any] + path: Path + + def __getitem__(self, name: str) -> np.ndarray: + return self.params[name] + + def linear(self, name: str) -> Tuple[np.ndarray, np.ndarray]: + return self.params[f"{name}.w"], self.params[f"{name}.b"] + + def num_parameters(self) -> int: + return int(sum(v.size for v in self.params.values())) + + def to_torch(self, dtype: Any = None) -> Dict[str, Any]: + """``{name: torch.Tensor}`` (float32 by default; float64 for the fp64 rewrites).""" + import torch + + dt = dtype or torch.float32 + return {k: torch.from_numpy(np.ascontiguousarray(v)).to(dt) for k, v in self.params.items()} + + +class _Reader: + """Node-addressed access to one ONNX file with the conventions of this export.""" + + def __init__(self, path: Path): + self.w = OnnxWeights(path) + self.facts: Dict[str, Any] = {"gelu": {}, "ln_eps": set()} + + def has_node(self, name: str) -> bool: + try: + self.w.node(name) + return True + except KeyError: + return False + + def const_input(self, node_name: str) -> np.ndarray: + """The single constant input of a binary node (the bias of a MatMul + Add pair).""" + node = self.w.node(node_name) + consts = [t for t in node.inputs if t and self.w.has(t)] + if len(consts) != 1: + raise ValueError(f"{node_name}: expected one constant input, found {len(consts)}") + return np.asarray(self.w.array(consts[0]), np.float32) + + def bias_after(self, matmul_name: str) -> np.ndarray: + add = self.w.consumer_of(self.w.node(matmul_name).outputs[0], "Add") + return self.const_input(add.name) + + def linear(self, path: str) -> Tuple[np.ndarray, np.ndarray]: + """``(w [in, out], b [out])`` of the torch ``nn.Linear`` exported under ````: either ``/MatMul`` + followed by an ``Add`` (3-D inputs) or ``/Gemm`` (2-D inputs); exactly one of the two must exist.""" + has_mm, has_gemm = self.has_node(f"{path}/MatMul"), self.has_node(f"{path}/Gemm") + if has_mm == has_gemm: + raise KeyError(f"{path}: expected exactly one of MatMul / Gemm, found {has_mm=} {has_gemm=}") + if has_mm: + w = np.asarray(self.w.matmul_weight(f"{path}/MatMul"), np.float32) + return w, self.bias_after(f"{path}/MatMul") + return self.gemm(f"{path}/Gemm") + + def gemm(self, node_name: str) -> Tuple[np.ndarray, np.ndarray]: + """``(w [in, out], b [out])`` of a ``Gemm`` node (``y = x @ W^T + b`` with ``transB = 1``).""" + g = self.w.gemm(node_name) + if g.trans_a or g.alpha != 1.0 or g.beta != 1.0 or g.bias is None: + raise ValueError(f"{node_name}: unexpected attributes {g.trans_a=} {g.alpha=} {g.beta=}") + w = np.asarray(g.weight, np.float32) + return (w.T if g.trans_b else w).copy(), np.asarray(g.bias, np.float32).reshape(-1) + + def layer_norm(self, path: str) -> Tuple[np.ndarray, np.ndarray]: + node = self.w.node(f"{path}/LayerNormalization") + self.facts["ln_eps"].add(float(node.attrs.get("epsilon", 1e-5))) + if int(node.attrs.get("axis", -1)) != -1: + raise ValueError(f"{path}: LayerNormalization over axis {node.attrs.get('axis')}") + return (np.asarray(self.w.param(node.name, 1), np.float32), + np.asarray(self.w.param(node.name, 2), np.float32)) + + def gelu(self, path: str) -> str: + node = self.w.node(f"{path}/Gelu") + mode = str(node.attrs.get("approximate", "none")) + self.facts["gelu"][path] = mode + return mode + + def scalar(self, node_name: str) -> float: + return float(np.asarray(self.const_input(node_name)).reshape(-1)[0]) + + +def _put_linear(params: Dict[str, np.ndarray], name: str, wb: Tuple[np.ndarray, np.ndarray]) -> None: + params[f"{name}.w"], params[f"{name}.b"] = wb + + +def _put_ln(params: Dict[str, np.ndarray], name: str, gb: Tuple[np.ndarray, np.ndarray]) -> None: + params[f"{name}.gamma"], params[f"{name}.beta"] = gb + + +def _mlp(r: _Reader, params: Dict[str, np.ndarray], onnx_path: str, name: str, gelu: str) -> None: + _put_linear(params, f"{name}.fc1", r.linear(f"{onnx_path}/fc1")) + _put_linear(params, f"{name}.fc2", r.linear(f"{onnx_path}/fc2")) + mode = r.gelu(f"{onnx_path}/act") + if mode != gelu: + raise ValueError(f"{onnx_path}: GELU approximate={mode!r}, expected {gelu!r}") + + +def _read_encoder(path: Path, params: Dict[str, np.ndarray]) -> Dict[str, Any]: + r = _Reader(path) + for cat, mod in {**MIXER_ENCODERS, **SMALL_ENCODERS}.items(): + P, N = f"/encoder/{mod}", f"encoder.{mod}" + _mlp(r, params, f"{P}/channel_pre_project", f"{N}.channel_pre_project", "none") + if cat in MIXER_ENCODERS: + _mlp(r, params, f"{P}/token_pre_project", f"{N}.token_pre_project", "none") + for i in range(C.MIXER_DEPTH): + B, BN = f"{P}/blocks.{i}", f"{N}.blocks.{i}" + _put_ln(params, f"{BN}.norm1", r.layer_norm(f"{B}/norm1")) + _mlp(r, params, f"{B}/tokens_mlp", f"{BN}.tokens_mlp", "none") + _put_ln(params, f"{BN}.norm2", r.layer_norm(f"{B}/norm2")) + _mlp(r, params, f"{B}/channels_mlp", f"{BN}.channels_mlp", "none") + _put_ln(params, f"{N}.norm", r.layer_norm(f"{P}/norm")) + _mlp(r, params, f"{P}/emb_project", f"{N}.emb_project", "none") + _put_linear(params, "encoder.neighbor_encoder.type_emb", r.linear("/encoder/neighbor_encoder/type_emb")) + for mod in ("lane_encoder", "route_encoder"): + _put_linear(params, f"encoder.{mod}.speed_limit_emb", r.linear(f"/encoder/{mod}/speed_limit_emb")) + _put_linear(params, f"encoder.{mod}.attribute_emb", r.linear(f"/encoder/{mod}/attribute_emb")) + unk = r.w.param(f"/encoder/{mod}/unknown_speed_emb/Gather", 0) + params[f"encoder.{mod}.unknown_speed_emb"] = np.asarray(unk, np.float32).reshape(-1) + _mlp(r, params, "/encoder/static_encoder/projection", "encoder.static_encoder.projection", "none") + _put_linear(params, "encoder.pos_emb", r.linear("/encoder/pos_emb")) + rpe = np.asarray(r.w.param("/encoder/Slice_5", 0), np.float32) + params["encoder.route_position_embedding"] = rpe.reshape(C.NUM_SEGMENTS_IN_ROUTE, C.HIDDEN_DIM) + scales, fills = set(), set() + for i in range(C.FUSION_DEPTH): + B, BN = f"/encoder/fusion/blocks.{i}", f"encoder.fusion.blocks.{i}" + _put_ln(params, f"{BN}.norm1", r.layer_norm(f"{B}/norm1")) + q_w = np.asarray(r.w.matmul_weight(f"{B}/attn/MatMul"), np.float32) + _put_linear(params, f"{BN}.attn.q", (q_w, r.bias_after(f"{B}/attn/MatMul"))) + kv_w = np.asarray(r.w.matmul_weight(f"{B}/attn/MatMul_1"), np.float32) + _put_linear(params, f"{BN}.attn.kv", (kv_w, r.bias_after(f"{B}/attn/MatMul_1"))) + _put_linear(params, f"{BN}.attn.out", r.gemm(f"{B}/attn/Gemm")) + _put_ln(params, f"{BN}.norm2", r.layer_norm(f"{B}/norm2")) + _mlp(r, params, f"{B}/mlp", f"{BN}.mlp", "none") + scales.add(r.scalar(f"{B}/attn/Mul_3")) + # the key-padding bias Where(mask, -inf, 0) is built once in block 0 and shared by the six blocks + fills.add(float(np.asarray(r.w.param("/encoder/fusion/blocks.0/attn/Where", 1)).reshape(-1)[0])) + _put_ln(params, "encoder.fusion.norm", r.layer_norm("/encoder/fusion/norm")) + return {"gelu": r.facts["gelu"], "ln_eps": r.facts["ln_eps"], "attn_scale": scales, "mask_fill": fills, + "sha256": r.w.sha256, "nodes": len(r.w.nodes())} + + +def _read_decoder(path: Path, params: Dict[str, np.ndarray]) -> Dict[str, Any]: + r = _Reader(path) + _mlp(r, params, "/dit/preproj", "decoder.dit.preproj", "none") + _mlp(r, params, "/dit/t_embedder", "decoder.dit.t_embedder", "none") + ego_row = np.asarray(r.w.param("/dit/Concat_3", 0), np.float32).reshape(-1) + expand = r.w.producer(r.w.node("/dit/Concat_3").inputs[1]) + nb_row = np.asarray(r.w.param(expand.name, 0), np.float32).reshape(-1) + params["decoder.dit.agent_embedding"] = np.stack([ego_row, nb_row]) + scales, fills = set(), set() + for i in range(C.DIT_DEPTH): + B, BN = f"/dit/blocks.{i}", f"decoder.dit.blocks.{i}" + _put_linear(params, f"{BN}.adaLN_modulation", r.linear(f"{B}/adaLN_modulation/adaLN_modulation.1")) + for n in ("norm1", "norm2", "norm3", "norm4"): + _put_ln(params, f"{BN}.{n}", r.layer_norm(f"{B}/{n}")) + qkv_w = np.asarray(r.w.matmul_weight(f"{B}/attn/MatMul"), np.float32) + _put_linear(params, f"{BN}.attn.qkv", (qkv_w, r.bias_after(f"{B}/attn/MatMul"))) + _put_linear(params, f"{BN}.attn.out", r.gemm(f"{B}/attn/Gemm")) + _mlp(r, params, f"{B}/mlp1", f"{BN}.mlp1", "tanh") + q_w = np.asarray(r.w.matmul_weight(f"{B}/cross_attn/MatMul"), np.float32) + _put_linear(params, f"{BN}.cross_attn.q", (q_w, r.bias_after(f"{B}/cross_attn/MatMul"))) + # K|V of this block: the bias is the constant input of cross_attn/Add_5, the weight a column block of the + # unnamed [256, 1536] MatMul whose output is split three ways (one 512-wide K|V slice per block) + add5 = r.w.node(f"{B}/cross_attn/Add_5") + kv_b = r.const_input(add5.name) + (kv_in,) = [t for t in add5.inputs if t and not r.w.has(t)] + split = r.w.producer(kv_in) + if split is None or split.op_type != "Split": + raise ValueError(f"{B}/cross_attn/Add_5: K|V does not come from a Split") + part = list(split.outputs).index(kv_in) + sizes = [int(s) for s in np.asarray(r.w.array(split.inputs[1])).reshape(-1)] + fused = r.w.producer(split.inputs[0]) + fused_w = np.asarray(r.w.matmul_weight(fused.name), np.float32) + start = int(sum(sizes[:part])) + _put_linear(params, f"{BN}.cross_attn.kv", (fused_w[:, start:start + sizes[part]].copy(), kv_b)) + _put_linear(params, f"{BN}.cross_attn.out", r.gemm(f"{B}/cross_attn/Gemm")) + _mlp(r, params, f"{B}/mlp2", f"{BN}.mlp2", "tanh") + scales.add(r.scalar(f"{B}/attn/Mul_2")) + scales.add(r.scalar(f"{B}/cross_attn/Mul_2")) + fills.add(float(np.asarray(r.w.param("/dit/blocks.0/attn/Where", 1)).reshape(-1)[0])) # shared by the blocks + F = "/dit/final_layer" + _put_linear(params, "decoder.dit.final_layer.adaLN_modulation", + r.linear(f"{F}/adaLN_modulation/adaLN_modulation.1")) + _put_ln(params, "decoder.dit.final_layer.norm_final", r.layer_norm(f"{F}/norm_final")) + _put_ln(params, "decoder.dit.final_layer.proj.0", r.layer_norm(f"{F}/proj/proj.0")) + _put_linear(params, "decoder.dit.final_layer.proj.1", r.linear(f"{F}/proj/proj.1")) + gelu = r.gelu(f"{F}/proj/proj.2") + if gelu != "tanh": + raise ValueError(f"{F}/proj/proj.2: GELU approximate={gelu!r}, expected 'tanh'") + _put_ln(params, "decoder.dit.final_layer.proj.3", r.layer_norm(f"{F}/proj/proj.3")) + _put_linear(params, "decoder.dit.final_layer.proj.4", r.linear(f"{F}/proj/proj.4")) + return {"gelu": r.facts["gelu"], "ln_eps": r.facts["ln_eps"], "attn_scale": scales, "mask_fill": fills, + "sha256": r.w.sha256, "nodes": len(r.w.nodes())} + + +def _read_turn(path: Path, params: Dict[str, np.ndarray]) -> Dict[str, Any]: + r = _Reader(path) + _put_linear(params, "decoder.turn_indicator_predictor", r.linear("/turn_indicator_predictor")) + return {"sha256": r.w.sha256, "nodes": len(r.w.nodes())} + + +EXPECTED_SHAPES = { # spot checks of the canonical layout (SPEC 4.2-4.4) + "encoder.neighbor_encoder.channel_pre_project.fc1.w": (C.NEIGHBOR_FEATURE_DIM, C.MIXER_CHANNELS), + "encoder.neighbor_encoder.token_pre_project.fc1.w": (C.INPUT_T + 1, C.MIXER_TOKENS), + "encoder.lane_encoder.token_pre_project.fc1.w": (C.POINTS_PER_SEGMENT, C.MIXER_TOKENS), + "encoder.polygon_encoder.channel_pre_project.fc1.w": (C.POLYGON_FEATURE_DIM, C.MIXER_CHANNELS), + "encoder.polygon_encoder.token_pre_project.fc1.w": (C.POINTS_PER_POLYGON, C.MIXER_TOKENS), + "encoder.line_string_encoder.channel_pre_project.fc1.w": (C.LINE_STRING_FEATURE_DIM, C.MIXER_CHANNELS), + "encoder.lane_encoder.attribute_emb.w": (C.LANE_ATTRIBUTE_DIM, C.MIXER_CHANNELS), + "encoder.turn_indicator_encoder.channel_pre_project.fc1.w": (C.TURN_INDICATOR_HISTORY, C.MIXER_CHANNELS), + "encoder.pos_emb.w": (C.POS_FEATURE_DIM, C.HIDDEN_DIM), + "encoder.fusion.blocks.0.attn.kv.w": (C.HIDDEN_DIM, 2 * C.HIDDEN_DIM), + "decoder.dit.preproj.fc1.w": (C.DIT_INPUT_DIM, 512), + "decoder.dit.t_embedder.fc1.w": (C.DIT_TIME_DIM, 512), + "decoder.dit.blocks.0.adaLN_modulation.w": (C.HIDDEN_DIM, 6 * C.HIDDEN_DIM), + "decoder.dit.blocks.0.attn.qkv.w": (C.HIDDEN_DIM, 3 * C.HIDDEN_DIM), + "decoder.dit.blocks.2.cross_attn.kv.w": (C.HIDDEN_DIM, 2 * C.HIDDEN_DIM), + "decoder.dit.final_layer.proj.4.w": (C.DIT_MLP_DIM, C.DIT_INPUT_DIM), + "decoder.turn_indicator_predictor.w": (2 * len(C.TURN_HEAD_STEPS) + C.HIDDEN_DIM, C.TURN_INDICATOR_OUTPUT_DIM), +} + + +def _check_facts(facts: Dict[str, Any]) -> None: + eps = facts["encoder"]["ln_eps"] | facts["decoder"]["ln_eps"] + if eps != {C.LN_EPS} and not all(math.isclose(e, C.LN_EPS, rel_tol=1e-6) for e in eps): + raise ValueError(f"LayerNorm epsilons {sorted(eps)}, expected {C.LN_EPS}") + scales = facts["encoder"]["attn_scale"] | facts["decoder"]["attn_scale"] + if len(scales) != 1 or not math.isclose(scales.pop(), 1.0 / math.sqrt(C.HEAD_DIM), rel_tol=1e-6): + raise ValueError(f"attention scales {facts['encoder']['attn_scale'] | facts['decoder']['attn_scale']}") + fills = facts["encoder"]["mask_fill"] | facts["decoder"]["mask_fill"] + if fills != {float("-inf")}: + raise ValueError(f"attention mask fill values {fills}, expected -inf") + + +def find_weights_dir(explicit: Optional[str] = None) -> Optional[Path]: + """A local directory holding the v5.0 files, or None: ``explicit`` > ``$DIFFUSION_PLANNER_WEIGHTS_DIR`` > the + workspace download (``assets/diffusion-planner/hf_diffusion_planner``) > the HF cache snapshot of the pinned + revision (``local_files_only``; never a network access).""" + import os + + names = (C.ENCODER_ONNX, C.DECODER_ONNX, C.TURN_INDICATOR_ONNX, C.PARAM_JSON) + cands: List[Path] = [] + for c in (explicit, os.environ.get("DIFFUSION_PLANNER_WEIGHTS_DIR")): + if c: + cands.append(Path(c).expanduser()) + here = Path(__file__).resolve() + for parent in here.parents: + cands.append(parent / "assets" / "diffusion-planner" / "hf_diffusion_planner") + for c in cands: + if all((c / n).is_file() for n in names): + return c + try: # the HF cache of `from_pretrained` (offline lookup only) + from huggingface_hub import snapshot_download + + p = Path(snapshot_download("AutowareFoundation/diffusion_planner", + revision="423efde67f5414734da43a7ad856c17ceb8b51aa", + allow_patterns=list(names), local_files_only=True)) + if all((p / n).is_file() for n in names): + return p + except Exception: # noqa: BLE001 -- not cached / no huggingface_hub: no weights + pass + return None + + +def load_weights(weights_dir: Path, *, verify_sha256: bool = True) -> PlannerWeights: + """Read the encoder / decoder / turn-indicator ONNX files and the param JSON of ``weights_dir``.""" + weights_dir = Path(weights_dir) + sha = {n: file_sha256(weights_dir / n) for n in C.FILE_SHA256} + if verify_sha256: + bad = {n: s for n, s in sha.items() if s != C.FILE_SHA256[n]} + if bad: + raise ValueError(f"{weights_dir}: files differ from AutowareFoundation/diffusion_planner@v5.0: " + f"{sorted(bad)} (pass verify_sha256=False to load another export)") + params: Dict[str, np.ndarray] = {} + facts = {"encoder": _read_encoder(weights_dir / C.ENCODER_ONNX, params), + "decoder": _read_decoder(weights_dir / C.DECODER_ONNX, params), + "turn": _read_turn(weights_dir / C.TURN_INDICATOR_ONNX, params)} + _check_facts(facts) + for name, shape in EXPECTED_SHAPES.items(): + if tuple(params[name].shape) != shape: + raise ValueError(f"{name}: shape {params[name].shape}, expected {shape}") + for name, arr in params.items(): + if arr.dtype != np.float32 or not np.isfinite(arr).all(): + raise ValueError(f"{name}: not finite float32") + norm = load_param_json(weights_dir / C.PARAM_JSON) + return PlannerWeights(params, norm, sha, facts, weights_dir) + + +def param_count(params: Mapping[str, np.ndarray]) -> int: + return int(sum(v.size for v in params.values())) + + +def coverage(weights: PlannerWeights) -> Dict[str, Any]: + """Proof that the canonical dict is a re-labelling of the export: every float initializer (size > 1) of the three + files equals one canonical tensor (or the concatenation it was split from: the fused cross K/V MatMul, the two + agent-embedding rows), and no canonical tensor holds the same data twice except the biases the export itself + deduplicated. Returns ``{"unused": [...], "duplicates": [...], "initializers": n}`` (tests require both lists + to be empty / the known pair).""" + import hashlib + + import onnx + from onnx import numpy_helper + + def h(a: np.ndarray) -> str: + a = np.ascontiguousarray(np.asarray(a, np.float32)) + return hashlib.sha1(a.tobytes() + str(a.shape).encode()).hexdigest() + + p = weights.params + known = {h(v): k for k, v in p.items()} + kv = np.concatenate([p[f"decoder.dit.blocks.{i}.cross_attn.kv.w"] for i in range(C.DIT_DEPTH)], axis=1) + known[h(kv)] = "decoder.dit.blocks.*.cross_attn.kv.w (fused)" + emb = p["decoder.dit.agent_embedding"] + known[h(emb[0:1])] = "decoder.dit.agent_embedding[0]" + known[h(emb[1:2])] = "decoder.dit.agent_embedding[1]" + rpe = p["encoder.route_position_embedding"] + known[h(rpe.reshape(1, *rpe.shape))] = "encoder.route_position_embedding" + for mod in ("lane_encoder", "route_encoder"): + unk = p[f"encoder.{mod}.unknown_speed_emb"] + known[h(unk.reshape(1, -1))] = f"encoder.{mod}.unknown_speed_emb" + for k, v in p.items(): # Gemm weights are stored [out, in] + if k.endswith(".w"): + known.setdefault(h(v.T), k) + unused, total = [], 0 + for f in (C.ENCODER_ONNX, C.DECODER_ONNX, C.TURN_INDICATOR_ONNX): + for init in onnx.load(str(weights.path / f)).graph.initializer: + a = numpy_helper.to_array(init) + if a.dtype != np.float32 or a.size <= 1: + continue + total += 1 + if h(a) not in known: + unused.append(f"{f}:{init.name}{list(a.shape)}") + seen: Dict[str, List[str]] = {} + for k, v in p.items(): + seen.setdefault(h(v), []).append(k) + dups = sorted(tuple(sorted(ks)) for ks in seen.values() if len(ks) > 1) + return {"unused": unused, "duplicates": dups, "initializers": total} diff --git a/code/tt_diffusion_planner/samples/README.md b/code/tt_diffusion_planner/samples/README.md new file mode 100644 index 0000000000000000000000000000000000000000..108db2ae3c18e4c5d5f0b019ea3bf299dfb51def --- /dev/null +++ b/code/tt_diffusion_planner/samples/README.md @@ -0,0 +1,23 @@ +# Sample inputs shipped with diffusion-planner-p150 + +Only REDISTRIBUTABLE data goes here (Apache-2.0 / MIT / CC-BY with attribution). Data whose license is unstated or +non-commercial (the Autoware demo rosbag, nuScenes, Argoverse 2, ...) is NOT shipped; the nuScenes-derived planner +instants used for local agreement tests live in the porting workspace only (`research/diffusion-planner/public_data`). + +Never use the `.bin` suffix here (or `.pt`, `.pth`, `.ckpt`, `.safetensors`): tt-model's staging silently drops +those suffixes from `code/` (`CODE_IGNORE`, tt-model-manager `src/tt_kernel/build.py:328-332`). + +Each sample is one `.npz` holding exactly the 15 raw (pre-normalization, ego-frame, batch-1, float32) tensors of the +node's `DiffusionPlannerCore::create_input_data()` (`reference/config.py` `INPUT_SCHEMA`): pass it as +`model(inputs=".npz")` or as the `/predict` field `inputs` (`server/client.py --inputs .npz`). + +Next to each sample, `.reference.json` is the `/predict` body of the fp32 CPU reference +(`tt_diffusion_planner.reference.ReferencePlanner`, `Output.to_dict()`, timing removed) on it; +`server/smoke_test.py` compares the served trajectory with it (ADE / FDE gates). Regenerate both the reference bodies +and `tests/goldens/` with `code/scripts/ref_golden.py` whenever the reference, the weights or the post-processing +changes. + +| file | content | source | license | +|---|---|---|---| +| `kashiwanoha_dense.npz` (default sample) | one planning instant on the kashiwanoha test map: ego at 6 m/s on a 17-lanelet route, 88 neighbours (constant-speed vehicles), 123 lanes (24 lanelets with traffic lights in the map), 60 line strings (stop lines and road borders), no intersection polygon; 113,893 B, sha256 `d8c2aaef...99f7` | Lanelet2 map [AutowareFoundation/map-carla-kashiwanoha](https://huggingface.co/datasets/AutowareFoundation/map-carla-kashiwanoha) 0.2.0 (`lanelet2_map.osm`, sha256 `4fe358f2...4e6d`, byte-identical to autoware_universe `planning/autoware_diffusion_planner/test_map`); scene scripted by `research/diffusion-planner/scripts/dp_scene.py` (`dp_reference.py --scene osm --osm .../lanelet2_map.osm --n-agents 120 --seed 7`), a Python port of the node's tensor construction (simplifications: SPEC 7) | Apache-2.0 (map: Apache-2.0 per its dataset card) | +| `straight_road.npz` | procedural 3-lane straight road: ego at 8 m/s, 12 neighbours (vehicles and pedestrians), 33 lanes with a green traffic light ahead, 6 route lanes, 1 intersection polygon, 9 line strings (borders + stop line); the precision-sensitive scene of SPEC 0.5; 10,832 B, sha256 `b9979b90...d381` | generated by `research/diffusion-planner/scripts/dp_scene.py` (`dp_reference.py --scene straight`), no external data | Apache-2.0 (this repo) | diff --git a/code/tt_diffusion_planner/samples/kashiwanoha_dense.npz b/code/tt_diffusion_planner/samples/kashiwanoha_dense.npz new file mode 100644 index 0000000000000000000000000000000000000000..6024fd09789439b7120fc5c70ed226ba534f9227 --- /dev/null +++ b/code/tt_diffusion_planner/samples/kashiwanoha_dense.npz @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:d8c2aaefae61e155f5281c8b31920820d442416c653b9354be33d1fe089399f7 +size 113893 diff --git a/code/tt_diffusion_planner/samples/kashiwanoha_dense.reference.json b/code/tt_diffusion_planner/samples/kashiwanoha_dense.reference.json new file mode 100644 index 0000000000000000000000000000000000000000..94cfc07c5b90dbed392711826e198179df09b5b7 --- /dev/null +++ b/code/tt_diffusion_planner/samples/kashiwanoha_dense.reference.json @@ -0,0 +1,963 @@ +{ + "model": "diffusion-planner-p150", + "frame_id": "base_link", + "meta": { + "predicted_agent_columns": [ + "x", + "y", + "yaw", + "cos", + "sin" + ], + "predicted_agent_rows": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11, + 12, + 13, + 14, + 15, + 16, + 17, + 18, + 19, + 20, + 21, + 22, + 23, + 24, + 25, + 26, + 27, + 28, + 29, + 30, + 31, + 32, + 33, + 34, + 35, + 36, + 37, + 38, + 39, + 40, + 41, + 42, + 43, + 44, + 45, + 46, + 47, + 48, + 49, + 50, + 51, + 52, + 53, + 54, + 55, + 56, + 57, + 58, + 59, + 60, + 61, + 62, + 63, + 64, + 65, + 66, + 67, + 68, + 69, + 70, + 71, + 72, + 73, + 74, + 75, + 76, + 77, + 78, + 79, + 80, + 81, + 82, + 83, + 84, + 85, + 86, + 87 + ], + "force_stop": false, + "time_from_start_s": [ + 0.1, + 0.2, + 0.3, + 0.4, + 0.5, + 0.6, + 0.7, + 0.8, + 0.9, + 1.0, + 1.1, + 1.2, + 1.3, + 1.4, + 1.5, + 1.6, + 1.7, + 1.8, + 1.9, + 2.0, + 2.1, + 2.2, + 2.3, + 2.4, + 2.5, + 2.6, + 2.7, + 2.8, + 2.9, + 3.0, + 3.1, + 3.2, + 3.3, + 3.4, + 3.5, + 3.6, + 3.7, + 3.8, + 3.9, + 4.0, + 4.1, + 4.2, + 4.3, + 4.4, + 4.5, + 4.6, + 4.7, + 4.8, + 4.9, + 5.0, + 5.1, + 5.2, + 5.3, + 5.4, + 5.5, + 5.6, + 5.7, + 5.8, + 5.9, + 6.0, + 6.1, + 6.2, + 6.3, + 6.4, + 6.5, + 6.6, + 6.7, + 6.8, + 6.9, + 7.0, + 7.1, + 7.2, + 7.3, + 7.4, + 7.5, + 7.6, + 7.7, + 7.8, + 7.9, + 8.0 + ], + "valid_counts": { + "ego": 1, + "neighbor": 88, + "static": 0, + "lane": 123, + "route": 17, + "polygon": 0, + "line_string": 60, + "goal": 1, + "ego_shape": 1, + "turn": 1 + }, + "reference": "fp32 CPU (torch)" + }, + "timing_ms": {}, + "num_poses": 80, + "columns": [ + "x", + "y", + "yaw", + "cos", + "sin", + "velocity", + "acceleration" + ], + "trajectory": [ + [ + 0.3642, + 0.0098, + -0.017, + 0.9963, + -0.0169, + 3.8048, + -0.0305 + ], + [ + 0.7626, + -0.0027, + -0.0413, + 0.9966, + -0.0413, + 3.8017, + -0.5598 + ], + [ + 1.1508, + -0.0284, + -0.0675, + 0.9941, + -0.0674, + 3.7457, + -0.4629 + ], + [ + 1.5358, + -0.0616, + -0.0936, + 0.991, + -0.0933, + 3.6994, + -0.5838 + ], + [ + 1.9181, + -0.1035, + -0.1189, + 0.9888, + -0.1184, + 3.6411, + -0.5708 + ], + [ + 2.2977, + -0.1548, + -0.1417, + 0.9862, + -0.141, + 3.584, + -0.6467 + ], + [ + 2.6624, + -0.2151, + -0.1675, + 0.9822, + -0.1664, + 3.5193, + -0.5376 + ], + [ + 3.0241, + -0.2833, + -0.1917, + 0.978, + -0.1902, + 3.4656, + -0.5445 + ], + [ + 3.3779, + -0.3594, + -0.2164, + 0.9725, + -0.2143, + 3.4111, + -0.5043 + ], + [ + 3.7213, + -0.4447, + -0.2416, + 0.9656, + -0.2386, + 3.3607, + -0.4697 + ], + [ + 4.0601, + -0.5401, + -0.2663, + 0.9581, + -0.2623, + 3.3137, + -0.4001 + ], + [ + 4.3847, + -0.6408, + -0.2939, + 0.9494, + -0.2885, + 3.2737, + -0.2939 + ], + [ + 4.7077, + -0.7432, + -0.3206, + 0.9408, + -0.3138, + 3.2443, + -0.1608 + ], + [ + 5.0188, + -0.8572, + -0.3486, + 0.9315, + -0.3401, + 3.2282, + -0.1422 + ], + [ + 5.3243, + -0.9728, + -0.3751, + 0.9216, + -0.3647, + 3.214, + -0.0311 + ], + [ + 5.6226, + -1.1007, + -0.4053, + 0.9108, + -0.3926, + 3.2109, + 0.0181 + ], + [ + 5.9167, + -1.2306, + -0.433, + 0.8995, + -0.4178, + 3.2127, + -0.0254 + ], + [ + 6.1979, + -1.3752, + -0.4623, + 0.8875, + -0.4442, + 3.2102, + 0.1051 + ], + [ + 6.4825, + -1.5215, + -0.4887, + 0.8753, + -0.4676, + 3.2207, + 0.0699 + ], + [ + 6.7579, + -1.6771, + -0.5155, + 0.8631, + -0.4912, + 3.2277, + 0.2097 + ], + [ + 7.0396, + -1.8412, + -0.5422, + 0.85, + -0.5142, + 3.2486, + 0.0579 + ], + [ + 7.3087, + -2.0142, + -0.5688, + 0.8368, + -0.5369, + 3.2544, + 0.1655 + ], + [ + 7.5771, + -2.196, + -0.593, + 0.8241, + -0.5573, + 3.271, + 0.2495 + ], + [ + 7.8456, + -2.3808, + -0.6156, + 0.8112, + -0.5758, + 3.2959, + 0.0594 + ], + [ + 8.1022, + -2.5712, + -0.6361, + 0.7995, + -0.5924, + 3.3019, + 0.3894 + ], + [ + 8.3573, + -2.7721, + -0.6565, + 0.788, + -0.6089, + 3.3408, + 0.3066 + ], + [ + 8.6075, + -2.9804, + -0.6752, + 0.7772, + -0.6238, + 3.3715, + 0.3165 + ], + [ + 8.8647, + -3.1919, + -0.692, + 0.7672, + -0.6371, + 3.4031, + 0.3109 + ], + [ + 9.1132, + -3.4101, + -0.7071, + 0.7577, + -0.6487, + 3.4342, + 0.4503 + ], + [ + 9.3631, + -3.6305, + -0.7217, + 0.7478, + -0.6596, + 3.4792, + 0.5671 + ], + [ + 9.6151, + -3.8648, + -0.7345, + 0.7392, + -0.6691, + 3.5359, + 0.3905 + ], + [ + 9.8532, + -4.0943, + -0.7461, + 0.7313, + -0.6776, + 3.575, + 0.5891 + ], + [ + 10.1062, + -4.3371, + -0.7559, + 0.7248, + -0.6848, + 3.6339, + 0.5732 + ], + [ + 10.3595, + -4.5775, + -0.7647, + 0.7185, + -0.6911, + 3.6912, + 0.4608 + ], + [ + 10.6115, + -4.8218, + -0.772, + 0.7126, + -0.696, + 3.7373, + 0.6758 + ], + [ + 10.8603, + -5.079, + -0.7786, + 0.7079, + -0.7007, + 3.8049, + 0.8014 + ], + [ + 11.1181, + -5.3397, + -0.7841, + 0.7034, + -0.7043, + 3.885, + 0.6251 + ], + [ + 11.3825, + -5.6107, + -0.7891, + 0.6994, + -0.7076, + 3.9475, + 0.553 + ], + [ + 11.647, + -5.877, + -0.7929, + 0.6961, + -0.71, + 4.0028, + 0.6381 + ], + [ + 11.9066, + -6.1516, + -0.7954, + 0.6942, + -0.7117, + 4.0666, + 0.7248 + ], + [ + 12.1832, + -6.4356, + -0.7983, + 0.6918, + -0.7136, + 4.1391, + 0.624 + ], + [ + 12.4523, + -6.7125, + -0.8003, + 0.6893, + -0.7145, + 4.2015, + 0.7835 + ], + [ + 12.737, + -7.0005, + -0.8003, + 0.689, + -0.7144, + 4.2799, + 0.7477 + ], + [ + 13.0302, + -7.304, + -0.8008, + 0.6885, + -0.7147, + 4.3546, + 0.4532 + ], + [ + 13.3201, + -7.6034, + -0.8006, + 0.6879, + -0.7142, + 4.4, + 0.6992 + ], + [ + 13.6126, + -7.9086, + -0.8013, + 0.6881, + -0.715, + 4.4699, + 0.5933 + ], + [ + 13.9119, + -8.2123, + -0.8003, + 0.6887, + -0.7143, + 4.5292, + 0.6839 + ], + [ + 14.2205, + -8.5202, + -0.7993, + 0.689, + -0.7134, + 4.5976, + 0.6703 + ], + [ + 14.532, + -8.8399, + -0.7965, + 0.6903, + -0.7112, + 4.6646, + 0.4373 + ], + [ + 14.8533, + -9.1531, + -0.7946, + 0.6917, + -0.7098, + 4.7084, + 0.6221 + ], + [ + 15.1759, + -9.4878, + -0.7917, + 0.694, + -0.708, + 4.7706, + 0.6287 + ], + [ + 15.4998, + -9.812, + -0.7876, + 0.6963, + -0.7048, + 4.8334, + 0.438 + ], + [ + 15.8367, + -10.1435, + -0.7863, + 0.6968, + -0.7037, + 4.8772, + 0.5713 + ], + [ + 16.165, + -10.4801, + -0.7826, + 0.6989, + -0.7009, + 4.9344, + 0.5701 + ], + [ + 16.5128, + -10.8125, + -0.7789, + 0.7016, + -0.6983, + 4.9914, + 0.4274 + ], + [ + 16.8627, + -11.1549, + -0.7766, + 0.7023, + -0.6963, + 5.0341, + 0.3822 + ], + [ + 17.2078, + -11.4905, + -0.7722, + 0.706, + -0.6935, + 5.0723, + 0.6469 + ], + [ + 17.5616, + -11.8417, + -0.7685, + 0.7072, + -0.6903, + 5.137, + 0.3521 + ], + [ + 17.9438, + -12.187, + -0.7655, + 0.7089, + -0.688, + 5.1722, + 0.3652 + ], + [ + 18.2915, + -12.537, + -0.7616, + 0.7111, + -0.685, + 5.2087, + 0.4565 + ], + [ + 18.6744, + -12.8863, + -0.7565, + 0.7139, + -0.681, + 5.2544, + 0.2074 + ], + [ + 19.0473, + -13.2427, + -0.7542, + 0.7151, + -0.6792, + 5.2751, + 0.4781 + ], + [ + 19.4273, + -13.5907, + -0.7495, + 0.7178, + -0.6757, + 5.3229, + 0.4162 + ], + [ + 19.8051, + -13.9481, + -0.7452, + 0.7201, + -0.6724, + 5.3646, + 0.4455 + ], + [ + 20.196, + -14.3107, + -0.7423, + 0.7217, + -0.6701, + 5.4091, + 0.2855 + ], + [ + 20.5865, + -14.6641, + -0.7398, + 0.7233, + -0.6681, + 5.4377, + 0.3681 + ], + [ + 20.9919, + -15.0273, + -0.7355, + 0.7255, + -0.6648, + 5.4745, + 0.1879 + ], + [ + 21.3863, + -15.3812, + -0.7322, + 0.7268, + -0.662, + 5.4933, + 0.4135 + ], + [ + 21.7792, + -15.7441, + -0.7282, + 0.7284, + -0.6587, + 5.5346, + 0.3407 + ], + [ + 22.1947, + -16.1108, + -0.7244, + 0.7304, + -0.6556, + 5.5687, + 0.2159 + ], + [ + 22.603, + -16.4771, + -0.7213, + 0.7312, + -0.6529, + 5.5903, + 0.2568 + ], + [ + 23.015, + -16.85, + -0.7174, + 0.7322, + -0.6495, + 5.6159, + 0.1228 + ], + [ + 23.4322, + -17.2175, + -0.714, + 0.7336, + -0.6466, + 5.6282, + 0.0 + ], + [ + 23.8479, + -17.587, + -0.7115, + 0.7335, + -0.6441, + 5.6282, + 0.0 + ], + [ + 24.2676, + -17.9567, + -0.7075, + 0.7338, + -0.6403, + 5.6282, + 0.0 + ], + [ + 24.6828, + -18.3369, + -0.7049, + 0.7344, + -0.6379, + 5.6282, + 0.0 + ], + [ + 25.1098, + -18.7025, + -0.702, + 0.7344, + -0.635, + 5.6282, + 0.0 + ], + [ + 25.5369, + -19.0821, + -0.6978, + 0.7343, + -0.6309, + 5.6282, + 0.0 + ], + [ + 25.9703, + -19.4509, + -0.6938, + 0.7344, + -0.627, + 5.6282, + 0.0 + ], + [ + 26.3879, + -19.8323, + -0.6905, + 0.7332, + -0.6234, + 5.6282, + 0.0 + ] + ], + "turn_indicator": { + "command": 1, + "command_name": "DISABLE", + "keep_selected": true, + "held": false, + "logits": [ + -15.99992561340332, + -5.138933181762695, + -5.255273818969727, + -0.7758127450942993, + 5.528584003448486 + ], + "probabilities": [ + 1.5498277106118508e-09, + 8.075185905909166e-05, + 7.188303425209597e-05, + 0.006339156534522772, + 0.9934089183807373 + ] + }, + "predicted_agents": { + "format": "npz", + "key": "predicted_agents", + "dtype": "float32", + "shape": [ + 88, + 80, + 5 + ], + "data": "UEsDBC0AAAAAAAAAIQBvdRbK//////////8UABQAcHJlZGljdGVkX2FnZW50cy5ucHkBABAAgCYCAAAAAACAJgIAAAAAAJNOVU1QWQEAdgB7J2Rlc2NyJzogJzxmNCcsICdmb3J0cmFuX29yZGVyJzogRmFsc2UsICdzaGFwZSc6ICg4OCwgODAsIDUpLCB9ICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAKUEr8v34bVEC/hKS/EdGXPo19eL/wz/i/umFRQEFOpL/g8pc+ARx4v4Dv9b/KpE5AAyKkv0nQmD5AFXi/4O/zv24cTED6/KO/T6iaPlZ6eL9osvG/0NJJQPxao78l5J0+x2V4vzA377+MaUdAzBijv0VooD7+zXi/0B/tv4oMRUA9ZqK//ueiPidPeL+Aguu/nh5DQNbvob+YfaU+Y1F4vwBc6b9kTUFAcEqhv5WBpz6rvXe/ULDnv8AsP0DbM6G/Ky+pPkgseL/w1uW/dTE9QGnGoL9Ydas+/SB4vzC85L8DRjtAWZygv9TYrD4pTHi/mM3iv/iEOUC9KaC/pV6vPtZLeL9gu+G/H/s3QKE5oL8K/a8+j6V4v/h74L/2YDZAYgWgvyrxsD7Skni/KH3fv9ehNEAdkJ+/zsCxPsztd7+gu92/bnQzQMEsn79YKLM+0KN3v0ir3L+fXjJAROOev9hztD5HhHe/iJTbv0odMUB9gJ6/thq2Pk1Rd78QVdq/ThcwQG77nb/YfLc+Wb92v8Cy2b8KDi9A0s+dv+g/uD5uq3a/eI3Yv4oOLkBCsJ2/dP64Pm2udr8oO9i/kO0sQPFnnb97ibk+3Et2v9A317+ETixAHXGdv36DuT55XHa/sNjWv7OzK0BfP52/PmO5PpTrdb/Q9da/7hkrQJ1Fnb8mJrk+w+J1v7ja1b9LVSpA9SSdvwDmuD6YiXW/oK/Vv3SAKUDWGZ2/l6u4Pl1edb8w2NW/GiopQMAjnb+ih7g+4mV1vyiQ1b/FWyhADOScv9jguD6lA3W/cGHVv2AhKEDjmJy/iAi5Ps94dL9omNa/1CQoQPx0nL9Oebk+a1d0vzCm1b/MdidAeiCcvw6auT5gt3O/kEPUv8lJJ0AZ2Ju/A566PlV/c7/QVNa/GA0nQHqFm7/oubo+qeFyv5jG1b/00iZAhxqbv5oTvD4pgXK/gPHWvyZwJkAX45q/Zka8Ppoicr94Pte/sl4mQLSHmr9WFb0+WrFxv0jX17/pXyZASvCZv7GBvj7v/HC/WMDXv36+JUCanJm/t1q/PnOecL9w4di/HMwlQMPzmL/5qsA+7rxvv/B52b9FmCVAAXSYv6rwwT5eKm+/4Fnav1qQJUAS25e/Ye7DPgCjbr9Y9tq/iKAlQF5Kl7+mRMU+PvNtv2hc27949iVAlaWWv6Csxj7dIG2/kFTbv2pQJUCJGZa/ER3IPqyCbL/wQ92/7dolQN16lb+x0Mk+C9Vrv7DY3r/t0SVAj8qUv8i/yj4Vw2q/MFPev6ihJUA8YJS/VNbLPv9Jar8gK9+/WYclQKC1k79uVs0+13JpvwCV378tDSZAWQqTv5Rpzz6oyWi/gArhv3YpJkAwg5K/s4/QPk4baL9QD+K/+DEmQHQKkr+pkdE+6n1nv5AJ47/mBiZA/42Rv/5q0z7THWe/YHDjv6QcJ0CNRJG/ys3TPkWrZr/w/+S/SQ8nQIz8kL/5ntQ+u15mv9BX5b84WCdAX7SQv35O1T4NB2a/OJLnv/HKJ0DIipC/byDWPvP2Zb9g8em/hmwoQGFbkL9ytdY+68dlv6BB67/XayhA50OQv+r21j72rWW/ONPrv5+iKEDhVJC/+6HXPscFZr8wqOy/GKAoQENwkL/uZNc+8Shmvxit7r+wmShAkHyQvz5h1z5BQGa/UK3vv7IlKUC8jZC/0c7XPiOFZr9I5fG/CdUoQIzakL9kRNc+TfJmv5Aj878DcSlAMQCRv4xO1z6ZQGe/GJD0v7QGKUBPPpG/IhPXPqmpZ79INva/U5UpQDBDkb/ERtc+3MNnvzgz+L/HXClAo1ORvxiR1z5r/Ge/OMD3v4iGKUAbS5G/rebXPqUGaL+AoPq/hXopQGlJkb8GvNc+qfVnv8Dx+78YkClAOS2Rvzga2D5T22e/COP9v/u6KUD/8pC/woPYPouIZ79Avf6/1LIoQGy7kL+PHNk+DUpnvyBL/b+4nyhAQlSQvyKP2T5uoGa/IIn/vxLXKEAMFJC/wGHZPiESZr8Qef6/iKsoQBemj78L99k+S2Zlv+gZ/b86wydA0nOPv7L02T6fAWW/EFX+v/rDJ0AJ2Y6/XyfaPlfeY78oqv2/WiEnQDJFjr8msto+4ORiv9INpcAeDEk/RveoPycQZT6BI3M/jH2iwGOsoD8l4aw/WMVkPsnYej/i357AmcbfPwtCsD9urmQ+0teAP4r4msC5OBBAfX+yPyzaYz4cEoM/quiWwEgsMkAJjrQ/77xgPqfkhD9c45LAikxUQFN/tT/MP10+dIOFP8RqjsDIUHdAg4i2P/RsWT5+MoY/oNuJwBKLjUAL2bY/GnpWPho2hj80xoXAmKyfQK8Rtz/WP1Q+PTSGP4YAgcA25LFAHFa3Pz2EUT6nMIY/pBp5wDFfxEDSd7c/rqBOPtADhj+0A3HAnoDWQHHutz+mjEk+TfOFPxRXaMA9CulAAT64P3x4RT7/1IU/3HNgwCEp+0A62rg/ASk/Pj/IhT/wC1nAtqoGQSCNuT8vWjg+QsSFP3yiUsBbqg9Bqmm6P7slMD4Gw4U//MlLwGiwGEF5Nbs/X4gqPs34hT/gMEXA5qghQVK0uz9COSM+6auFP2hWP8DbfipBtSS8P9iJHD6+YIU/+Eo5wFRcM0FJubw/Mo0VPqwxhT/IFzTAiB88QaA+vT8IEw0+A8aEP1w/LsDtv0RBEqq9P5qSBT6IWoQ/oOApwKhbTUFzeb4/hnb2PdMBhD8IKyXAnLZVQUQLvz+FhOU9Vp6DP3gEIMClQF5BWze/P82b1z3K/YI/eJ8awHK8ZkE6lr8/WDXIPX16gj+4nRbAHfduQQfYvz+zPb89TziCP7weE8AKGndBZu+/P/FDsT3yf4E/oE8PwIYcf0Gv/r8/NjWpPVsXgT+4fgzAVpiDQU47wD/etp898MaAP4ADCcC0l4dBtTfAP+DsmT3gbIA/xKsGwHqFi0HFOMA/CtOSPeoDgD9gGQPAuFGPQRQGwD+MrY49oyZ/P1gxAcDzKZNBYjTAP6Q4iD1wwn4/oKX9vwUIl0EULsA/DlmFPRVgfj+Q0Pe/79KaQfEzwD++r4I9TBx+P1A/9b/8iZ5B+VrAP0Bbfj0jAX4/yODvv+o0okHEKMA/oKx7PVJ1fT9AE+u/Qb6lQZ2PwD/G/XA906F9P+gY6r9tV6lBeQnBP/4VZD1S0n0/iInmv7rGrEGcHsE/FhZdPROTfT/gzuK/gG2wQfR/wT/eyFM9Rsh9P2i037/AqLNBY8rBP5alRz3MpH0/GEnfv/gVt0EvO8I/1Kw7PXPPfT+Iftq/NXO6QWHhwj8WKDE96Hl+PxAK17/EvL1BRfjCP8ZTKj1KP34/+HTVv8oCwUFrhsM/ajcnPeMqfz/4wdG/tS/EQbTUwz9gvR09uDV/P3BXz78yZ8dBkfnDP/RJGD2AK38/CMPKvxChykGIs8Q/boIOPScEgD8w+sS/x5vNQZiqxD8S8xM9QyWAPzDSw79wvNBBvOvEP6jACT2ZF4A/ID2/v5jd00FUmMQ/Ym4VPWoegD/og72/Mv3WQfHSxD/anxQ97VKAP2Cyur/q89lBYZnEP3ZLHD1/VIA/mDG4vwHr3EFYUMQ/TqIhPXc0gD94Ure/uszfQYdKxD+kyCc9CF6AP/DXsb9QxOJB+7HDP4wvNj3EM4A/SKivv66h5UG8VcM/pKk/PeofgD8gHK2/nnvoQaitwj9Qu0w9w7Z/P6Dspb+yYetBOgPCP0aBWj1BM38/6Oamv5Zb7kFYk8E/wBdmPRQDfz/oZKG/GiDxQcaowD+UOno9Nl5+P0Bznr8r1vNB2hnAP5hWej3KQ30/YMycvxyR9kH6M78/uIGFPVR0fD+wdZe/TVj5QfR5vj8WD4k98W17P7izlr8G4/tBToq9PxzujT2eJXo/KAySvxyN/kGHmbw/lOeUPS8ZeT/4hZC/744AQsJluz9ylpk9mEh3PxCii7964gFClJG6P88tnT0fFHY/SMOMv/ISA0Kvdbk/CvegPShcdD8Ih4m/RlwEQv9HuD9EgaQ9lH1yP+ilgr9LmQVCb3K3P8s1qT3PaXE/oCOBv2zcBkJeQLY/wkqwPfLobz8Q8Xe/0iwIQicitT8BIbU9EFFuP6A9bb/jRAlCqsWzP2EKvD1jfmw/IEBnv6NiCkIagLI/jVXBPeasaj+gC2S//6MLQqNJsT9fh8Y96vZoPyAWXb9SpAxClj+wP+WbzD31q2c/wGtRv0nRDUI7Nq8/Jq/XPVTlZj9q1wfB9TqEwNZ+OD/wrDY/VkolP/Z3AMH3mHbAaAk5P26iOT9O8CY/yL3xwK76Y8A22jk/vmE6P/AGKD8IQeLAGrFPwAr0OD/6ADs/r18nP0w308BY4DrA1LA3P4eqOz/8XyY/YBHEwJzSJcBMrDY/gok8P8ixJT9wu7TAmQYRwNyfNT+zuTw/Y7okP9ygpcBoe/e/7bc0P0BlPT+jFCQ/UNCWwNmnzb9ytzM/zKk9P4YwIz+mY4fAW4Glv7fIMj+FJj4/bnIiP+DEccD+XXG/n+8xP7UcPj/BmCE/4CZUwDqVI7/uqTA/BSk+P3RcID/Q6zfASBWxvk9WLj/u0j8/OqoeP0hgGsDoQl29WKwqP3QXQT+JgRs/aA3/v6FVZD5eXSc/n89CP6LYGD8QWMm/pCb5Pqj/Iz/wMkU/WlcWPyA+lb9XST8/s0ggP/elRz8sfhM/wB0/v8vHfT9TOxs/cAlLP+yZDz8Aybe+0ECbP1sRFz8uhU4/W5IMPwCQKD3ne7Y/8K0SPxbeUj93iAk/QLPcPnzB0D9DGw4/GiRXPw07Bj+wL1E/zIjnP+ARCT/HUls/e2ECP2C8mD9q+/4/JcgDPwArXz8DQvw++CLIP2xxCUCx/vs+MiBjP++z8j7At/g/JjITQBJu7z6xEmc/UgToPiijFEAMQxtAoG3hPkvkaT9fP9s+IBksQOONIkAxs9E+Rp9tPx8MzT6S0EJAs1MoQN1cwT48+W8/l5G9PtKsWEDYXy1AwGSvPmzucj+km6w+UlNwQKTkMEDqi5s+vut0P69cmT4SUYNAfxs0QJRAhz6ETnc/AryFPmm+jUCbVDVA0dJhPlt/eD+wa18+AJaYQDGeNkAgpjQ+PvZ5P6L7Mj4JiqNAD5o2QNnyAz7W3no/x84CPmQmrUCjZTVAEfmlPWcNfD+OyKQ9hy23QPokNECc+wc9WYJ8P3ARBz2UYMBA3/EwQOgmbbwCBn0/BMdrvLANykBLnS1AlpeDvchWfT8J9IK9+XPTQCb6KEAD8OO9QbZ8PzK14r3Q0ttAmeEjQK2RHr5JWHw/ncAdvuMn5EDCWB5AWWdNvuxtez9iQUy+uZrsQEpyF0DqxHm+U9B6P+p4eL4NH/VA/rQQQEeakL5223k/N9OPvgyM+0BPmAhAXa6jviasdz+HZ6K+ONYBQXPbAECa7rS+5fd2Pw2ds77zswRB4V7xP4hVxb5zlXQ/wlvDvoQzCEG3ruA/N+/TvsYQcj9bKdG+k+AKQUYKzj9SmeC+KNNvP0YP3b4FPA5Bxt+9P+jd676tEm4/irjnvpAvEUE9K6w/7Ur1vktyaz8RDvC+7hIUQUzYnD+EKf6+REtpPysB+L4mbBZBiJuMP26NAr+Y2WY/wM39vl7kGEEds3s/a30Fv83sZD8vYAG/HgYbQdkQWj8kjgi/97FiP6LhA79v5xxBUN4+PyFbCr/K+GA/7DoFv1kEH0HC2CQ/C5YMvxGnXj/X1Qa//UMhQSwcBz+ngw2/vxJdPwlTB7+jLyNBDfDfPtPjDr9fQls/yTAIv86sJEFYfLM+10YPvxw9Wj8vSQi/f2gmQcjuhD5woRC/IK5YPzIxCb+8cyhBLCA/PqIkEb8Z0Fc/4HMJv6lvKUFMZ9M9bzQRv76lVj+2LAm/Xi8rQez+Ez2nNhG/GDpWP4wPCb/sZSxBGI9xvcyVEb+0slQ//PsIv2oxLkGciN297ccRv3JkVD/lFgm/8ngvQfYtIr7NJhG/tzRTP28eCL9exzBBbTNNvkR6Eb9z2FE/XgsIv5xZMkHaWIW+cpwRv9VdUT9HCQi/LHwzQYRon74dkRG/DCdPP3hYB7+ZRDRBj6uwvkVcEb9a8U0/OsoGvzhxNUE/Jb2+mGERvxeQTD9YaAa/VjA2QTS41L4tfhC/qBpLPy8eBb/suTdBnjjqvmY1EL94bUk/RFsEvxznOEFaovu+FNAOv294SD9muwK/5/I4QWkg/r7vfw6/zl5GP3vUAb96IzpBTegNvwGkDb+NtUQ/u4gAv3LvOkEJOBO/U3gMvzfiQz8yXv6+a4Y7QRFQFL/S0wu/ojhCP4Q2/L4ebDxB2Fcev6aYCr98WUE/TGX5vtbyPEFuYyS/cw8Jv4ZPQD+z7PW+0IcUwZbjGsHHiQ7Af/AcvywpS79U/xrBpTYhwV7tDsCayh6/LYJKv9hoIcF0hyfBxRIPwOA8IL9roEq/XqYnwdfiLcE9JA/ArU4hv0jfSr8IDi7B3Fc0wbYPD8BAjyG/flFLv7QFNMGD7jrB2S8PwLL4Ib8VA0u/cgw6wTVxQcGwHw/AxeIhv6Q5S79M8z/Bp/RHwUJZD8C7ISK/7G9Kv6LORcHqjE7BrGoPwPkXI7/HoEq/eJlLwU4PVcG1kQ/Auaojv+pJSr/UKFHBeaVbwZzPD8ANkyS/er9Jv/ypVsFWHWLBjBoQwDmsJb8DF0m/wttbwVSlaMEpXBDAEA8mv9w8SL+IW2HB5ABvwW+qEMCssia/LU5Hv3ySZsEIX3XBBOUQwAEIJr/PEUa/fLtrwSTYe8G1/RDAIdwlv/OZRb9cnXDB/PeAwX8rEcCIGyW/KohEv9qndcHcCoTB+VMRwCh1JL+xmEO/tpR6wUIah8H/hhHA/Bskv6WjQr8qjX/BiiSKwYCiEcADDSS/Hy9Cv+hOgsF/Go3Bt8sRwF9ZJL8+rkG/J7iEwQwUkMEN7RHAxsMkv3paQb+BFIfBzvSSwcUmEsD6gyW/xsxAv0OEicGF35XBnTESwMZgJr+DBkG/GPuLwe2pmMG6dxLALqInvyeBQL+8QI7BzGKbwRyxEsDUryi/KhZAv066kMHCL57BzMoSwPxJKr+saEC/6OOSwRHZoMG62xLAT2grv+ClQL9GRJXBHoejwbH8EsDUhSy/2KFAv9B5l8H0E6bBzgwTwMyBLb8h0kC/mNGZwXLAqMEcFBPAwEUuv5oMQb8GG5zBBUSrwaccE8DpyC6/5yRBv4NRnsHqyq3BYicTwPpsL78JQ0G/S3igwQRAsME0JRPA8aUvv2JlQb9S0qLBsaCywTYvE8DzQDC/QYJBv8KrpMFcELXBOjQTwOBeML9Re0G/xvGmwSZot8G9UBPApqIwvwUmQb9VFqnB1cC5wRNxE8BBpDC/tqNAv7xGq8FJI7zBxJITwIPVML8zMUC/9jOtwRpVvsHgyBPA+/YwvxplP7/GXq/B3pPAwdHqE8Bs4jC/3tI+vzZYscFxysLBiSQUwCVcMb9PHz6/QkSzwQrtxMGtYBTAxgsxv6AJPb/mVrXB/x7HwUKaFMBHLDG/+y88v7BMt8HvN8nB5/8UwOKqMb8rzjq/nB65weRLy8HqQBXA2yYyv13+Ob+o/LrBPlfNwVaFFcDvHDK/t+c4v+gEvcF2ac/B3dMVwPpoMr+RzTe/ROe+wZqi0cFUCBbAddwyv2ssN7/g5MDBinzTwa1KFsCzFDO/zzo2vxyTwsEEZNXBF3cWwNggM7+HjjW/mmbEwR1w18EzjBbAyvIyvx4nNb/ORMbBtGbZwcCSFsCkSDO/vDA1v4AtyMHyNNvBFp8WwL8gM7/w7jS/uubJwa8S3cG3nBbAXEwzv4cKNb8wusvBUhnfwUGqFsDj/TK/7LM0vzhzzcF0DeHB3YUWwOBqMr+8BzW/4kXPwcH34sHFRhbAj+Axv0/JNb/I2NDB7tHkwR4FFsAFezG/W6Q2v4Kr0sErxebB4MIVwJDKML/3YTe/CijUwQ+o6MHscRXAk8gvvxo3OL9M+tXBHqPqweMRFcD+XS+/q4g5v6Sm18E4p+zBWpwUwEk4Lr8x3zq/FFrZwS1/7sG5PBTAhD4tv2LwO7+WOtvBj4/wwRnEE8Chiiy/qIM9v4ri3MHYcfLBhUgTwL4wK7+f2D6/5FTewdJe9ME54BLAwLcqv0FEQL/QDODBSmH2wYxwEsAhNim/flVBvwzT4cG3XvjB3f0RwOhQKL/DuEK/pmbjwXpy+sGgphHAum0nv02uQ7/aJOXBlHn8wXE0EcD4lia/RhVFv1Se5sFZq/7BHNIQwIccJr+LZ0a/2nDowQRYAMJmfxDAPAYlv0sxR7/QMerBA1QBwrs8EMAV/CO/Vr9Hv7i868GDXQLCMQQQwKe0I7/fgEi/2o7twQxeA8IR2g/Ane4iv7bLSL/+9e7B2nQEwmGYD8B8ziG/iElJv/7V8MEMZgXCKqcPwEhaIb8q1ki/6HPywVZuBsLDeQ/APoYgvzMmSb/wRPTBgIEHwiB1D8AyVCC/oSBJv5Ewo0E8oXLBObY+v3AXOj9RvCy/FnGkQQKydMGIfj6/BBE4P5e6K79HraVB+ep2wVHqP78pIzY/+WAsv8zrpkGQK3nBoo5Bv4jIND8xdy2/WjKoQQxwe8FLr0K/lAo0P51ILr9/fqlBYc19weQjQ7+IsjM/iZguvyrJqkGtE4DB+jZDv7LZMz8Wuy6/eBesQaM7gcGu1EK/d4A0P7ScLr+RXK1Bs2OCwV8QQr+2ijU/uEQuv3ytrkFTjIPBBxZBv6bTNj9Rzy2/KvavQRDShMGyYUC/1is4P2KkLb8PNbFBCf6FweYxP7/iDzk/MNAsvy53skGbIofBGso9vwZXOj8w6iu/NrCzQf1IiMExKjy/MAE7P5yOKr+B+bRBCnOJwWrJOr8Sojs/720pv7AttkHJporBIhk5v5s1PD8R+Se/7XC3Qc3Ai8FfsDe/D6c8Pzy+Jr/it7hBLeWMwUJHNr+KUD0/6Zclvzb9uUFnCY7Bl3U1vwD0PT/xBCW/Yze7QWIwj8HRnzS/woQ+P61mJL9ehrxB+FKQwWGFNL95uz4/oWAkv6e/vUHCf5HBEPszv3JRPz+pDiS/1BO/Qe+qksF0SzS/Yq0/PyWAJL8jacBBcNyTwTJpNL9/HEA/hcYkv3i8wUFa9JTBVT00vx2oQD87ziS/Cf3CQdIilsF2pjS/seVAP35NJb+wVcRBllKXwQsgNb/wO0E/b+YlvzGvxUF7gZjBfNU1v1pRQT9hoya/yAHHQdi1mcFvODa/PpdBPyogJ7+7VchBq96awSWjNr/wqUE/0ZEnv+m7yUEsG5zB+zo3v3LgQT8jPii/xP3KQcQ8ncHl9je/ZolBP7/ZKL/GZ8xBaHiewf+nOL+5sUE/e5opv1rczUHWrZ/BhOE4v7bUQT9o4Sm/fCzPQUndoMFz5Dm/JH1BP0XEKr8amNBBhyKiwQgoOr/IhEE/FgsrvwTn0UGaX6PBHcA6v+t1QT9Vniu/elrTQauppMEWpDq/nbFBP+yYK7/ButRB/uClwSvJOr+jakE/IqMrvz0s1kG+JKfBDpQ6v46BQT92diu/Y5jXQeBqqMGbyzq/plBBP6mbK7/JBtlBJLqpwa3vOr+APEE/O7grv3tf2kHk96rBdgY6v6l/QT9n5yq/Zc/bQS9NrMFf4Tm/Ch5BPw2dKr/CQ91BpJutwepcOb9O3EA/QP8pv8e+3kFH+67BhZ45vxI1QD+YASq/TgfgQT9HsMFKODm/oAZAP66JKb9FneFBVZqxwYUcOb8AVz8/fSspv7cB40GlD7PB2yQ5v26qPj+K8ii/92LkQcZhtMHs+Ti/weI9P0h8KL+W1OVBA6C1wUGrOL9cgj0/pgkovxdM50ENDrfBnis5v2TxPD9wUii/T8XoQR9puMHKozm/R8o7P+hZKL/wIOpBsdC5wU0FOr9CFzs/jnYov6Kw60FsE7vBInI6v0lmOj/vnii/bhntQcqLvMHxQDu/bX85P04TKb/KgO5BzPe9wfETPL9M3Dg/SqUpvz4A8EFwdL/BEFQ9vyY0OD/eoCq/q3DxQaDNwMHIGT6/Ic83Py89K7/L2PJBkkXCwcRUP78QPDc/Zjssv11t9EHRxMPBfHFAv7JGNz+lWS2/g9r1QcJHxcElwEG/TJQ2P+VeLr/AVvdBq9PGwc0WQ78bajY/Q6Ivv8bD+EE6TcjBeUBEvzYWNj+AqDC/BEf6QWrsycEj30W/6bY1P74eMr+uwvtBd17LwbzaRr+3uDU/ZBozv8Ak/UGk78zBEA9Iv1qyNT/LSzS/MbH+QV+BzsE6JUm/Poo1P5BRNb+yBQBCDBbQwWMESr9qbzU/ASY2vyrRAEIusdHBDuFKv2RfNT+t/Da/u4wBQtJa08F3LEy/nv40P+ogOL+tRQJCuf7Uwd4fTb9ypjQ/Y/A4vzwDA0I2pdbBPhJOv9gWND87pzm/sM0DQpFe2MHu1U6/MOAzP99UOr+hjARCHRbawXNqT7/WOzM/A6Q6vzI5BUL8vdvBa0BQv99bMj8FGzu/bBAGQoCH3cEI4FC/bMAxP3R4O7/uxAZCCDDfwTlkUb9B+zA/+Kc7v1BzB0KizODBwNVRv2jILz+/lDu/ejEIQmSO4sFwgFK/0cwuP0TSO7+QoZjB09ylwUF3FMCAIyS/Ag83v9/pm8GJ16jBzscUwBtgJb/FXTa/NCmfwdG/q8FjmBTAi6Amv3eiN79yZ6LBzKGuwdYgFMCMrye/Ze85v4/QpcHylLHBna8TwI48KL997Tu/FQupwWSdtMEwVxPAuRopv6GvPb/GbqzB8pi3wXECE8C3sSm/WUU/v7ytr8GXjrrBORETwC6xKr/rfD+/QgGzwReUvcEyFBPAV8srv5zvP79qSbbBu3jAwZJmE8AbAC2/2i4/v1qAucGlfMPB8hgUwI7OLr8ZLz2/jLS8wXpFxsEf+BTAkU8xvxDGOr+E2r/BrinJwZ0TFsDGgDO/b0Q3v1Yxw8Ec68vBKTEXwL55Nr91CTS/4lrGwYK0zsFLkBjA46w4vyJyL7/ofMnBMmrRwUnMGcCGFDu/2ngrv36MzMGa+tPByiAbwCP/PL8L6ya/RKXPwfiT1sFhRRzAcrQ+vysGI79s3dLBjhbZwbdmHcDeT0C/oSMfvyr+1cEokdvB9FIewANAQr9pLhy/BEnZwYP83cFfLh/A7XxEvyuSGb+6adzBlEjgwXnxH8AycEa/tjgXv7qx38EeoeLBe74gwNyrSL82zBS/KOHiwQfm5MGoVCHAByNLv95GE788JObBNP/mwV8LIsDqCE6/v14Rv2ZO6cE3DOnBy8giwFB2UL87MA+/WrDswVon68FwUyPAY3tTv8v0Db9e1e/BwR3twXriI8C891W/LnkMv2Qp88GF+u7BCn0kwIeHWL9+0Aq/oED2wRay8MEwPiXAIAVbv32CCL9SsPnBFIHywefMJcDeDV2/CtkGv8YJ/cEoJvTBIoEmwOoqX79/mgS/sSIAwqHI9cGIMifAwSdhv3RaAr8xvAHC9Fj3wd30J8DDDGO/0Zf/viV2A8KLz/jBBb0owBWQZL/WEfq+pf0Ewocf+sHDrynAbYVmvyNo876etAbCGmv7wWmBKsBtDmi/pYztvlxLCMLYsfzBSpYrwA5+ab8MgeW+WvYJwkjs/cHemCzA9rJqvyvq3b6HfgvCCuL+wUDRLcBjO2y/lsPUviI1DcL53v/B0vEuwJsobb8GF8y+88kOwhttAMJrPzDAR3RuvyUmwr7ERhDCm9MAwtWPMcDEJW+/JOK3vt76EcIzKQHCD+0ywK8IcL/RS62+YnUTwrByAcJ5pDTAfC5xv778n77IBBXCp7QBwqQlNsCMXnK/TF+UvriLFsKk4gHCy9M3wEg6c797QYe+8ygYwlUaAsKDpznAELtzv2TIcb7Q0hnCoEwCwtNOO8DuFXW/wjRYvhBoG8JpWgLC3iY9wJvkdb9GVzu+QuIcwhNKAsKMAD/AjOF1v+YZHr7HWB7Cl1ACwv3mQMDcT3a/8kEAvlDpH8L4UwLCib5CwMOcdr+Knsa9N4chwngmAsLSoETAFMJ2v8Nfi71w5yLC8wMCwmtlRsBWFHe/FKknvWRyJMJt6gHCYWtIwGXWdr8diCG8jPMlwn22AcIL4kdAY1N2v/sOlDwvayfCL4ABwocaRkDMA3a/Hqw5PSLgKMJKMQHCcHdEQGvOdb+rL5A9EnIqwp7hAMKMwEJAbL11vyMHxj1bwivCi4YAwi8fQUAsHXW/QgD5Pc5GLcL8LwDCt4Q/QB4Adb92sRU+Ob4uwrKo/8FxJD5AWw51vwxlKz7IKTDC8Yv+wZymPECg03S/79pCPoGeMcKkvf3BvFQ7QLjwdL8bx1c+JBMzwkOq/MHnGzpAl6J0v6wBaz4WbDTCGID7wfT4OEDv+3S/Qzp9PhO9NcK5X/rBato3QPJ7dL/qY4c+VBw3wiYs+cEC0TZAdIp0v+qrjz7mlzjCzAT4wTXhNUAekHS/vimXPgvkOcLi1fbBcCM1QPNWdL89B50+sC47wrmG9cHtYTRA7Tt0v/sNoz7NgzzCKUX0wZmvM0BSyHO/J3+oPkTmPcLu1fLBzuYyQP4xc7/SmK4+1TI/wv6B8cGzSDJAzIByvxlRsz7zbkDCbtjvwY3DMUDOZ3G/IBm3Phi6QcLpkO7BOVMxQO1ScL9hOLo+0hVDwubv7MHHyDBA5hlvv4gYvj4HW0TCunPrwYppMEAKlG2/03zAPqmgRcJo8unB8hYwQCGNbL9xqsI+cljYwWKI4kHrmke/UkovP9Y1Mb+IFNXBM1ffQQgjSr8LAjA/0AE0v6DG0cGcGdxBgwxOv5fZLz+B0ze/dHfOwaDX2EHwpFC/s1kvP7EzOr+QRcvBBoPVQZ9SU7/UKi8/DM48vwj2x8HoHNJB4DxVv6m9Lj+Vij6/vrHEwdSuzkEIfVe/Vr0uPyPPQL90gcHB8kTLQR4vWL8oBy8/QKRBv0BMvsFG18dBg+9Yv1KXLz8PqEK/viC7we53xEGW/Fi/rGowP6sUQ7/q87fBZxnBQbVnWL/ELTE/K9VCv6zetMHRv71BhplYvy+SMT8PNUO/rNSxwW5rukHHoFe/KDwyPxOEQr863K7BlTa3QSTGVr/EDTI/6pBBv1Hbq8F4ErRB0OxVvz1gMj/P2EC/Rv6owQTgsEEFelS/QUIyP31TP7/SH6bBntCtQdJjU7+2IjI/LSw+v0xIo8EW06pBTYtRv+6TMj9+gDy/92ugwdTXp0FQYlC/UrwyP7FmO7/utp3Bt+2kQVYyT7/RuDI/bjM6v+z+msHEFqJBb1BOv0qWMj/iQTm/wlSYwaNQn0HxzUy/1NIyP0zYN7/UwpXB+oOcQRzSS79S1jI/3t02v9P9ksFH0plBKHpKv3hkMz/OwTW/aZOQwXcsl0FWIEm/wc0zP3yUNL8mFo7B5I2UQSzyR78xazQ/Oqgzv+eLi8FYF5JBCq5GvzjzND8OnTK/5TCJwbyjj0H1uUW/Vnw1P+LhMb+c34bB5DuNQcbCRL+bGTY/iysxv5aAhMFr3opBeRZEv/IoNj9FhjC/NVGCwbiUiEFh2UK/2qg2P1d+L78wOYDBemSGQSptQr/54DY/OSkvvxDGe8FqOIRBqcdBvyY0Nz/qpS6/QH13wd84gkEGvEC/LqQ3P4fILb9oaXPBOyyAQbuzQL8XZjc/wactv25Ob8GVHHxBE5lAv0meNz+Boy2/AFVrwQ1PeEFbEEC/R+c3P5k4Lb8AhmfBxKJ0QYphP7+NDzg//5osvyybY8Hdy3BBwBE/vwTwNz+KPyy/hK1fwV5RbUGBzD6/SiA4P80NLL+kzlvBE9BpQe0ePr/SPzg/KG4rvyKKWMFfpWZBijI+v3v8Nz9YZyu/PrdUwTxlY0EC9D2/21Y4P6BMK7+oTVHBLTlgQSPfPb8XrDc/j/Uqv14HTsGHI11B2Vg9v8WLNz9UZCq/iJxKwWIDWkF7dD2/y5I3P1iCKr+Ej0fBgC5XQTLSPb8v+zY/AKQqv4CARMFiVlRBN+I9v3dnNj9eeiq/ID1BwYZuUUGlGD6/oPE1PyaCKr+iPT7B2NtOQWkePr+SszU/q28qv3SPO8FSgkxBYDA+vwUQNT+lQSq/jK84wRMUSkEMez6/NpY0P4tbKr9qODXBVm9HQfoFP7844jM/p50qv37iMsGMN0VBUJA/v8VmMz8N9Sq/4vsvwRsvQ0Epnz+/p78yP/zBKr/QSS3BauRAQb0IQL9JWTI/MgErv1apKsG0fj5BfolAvyjEMT+FRCu/1PgnwTM9PEEhR0G/tmIxP7HXK7/cNCXBV1I6QQ/LQb9sJTE/mkAsv2TxIsHiFjhB07NCv5LoMD9/DC2/yP4fwa0eNkEKQ0O/I+4wPzubLb8qrx3BV+UzQZg1RL8GczA/2lcuvwxkG8ECzTFBUeFEv6g/MD/F6y6/FvgYweHvL0Gp70W/zB0wP9HnL7/KfhbBRrotQSRCR7+zqy8/YwYxv34ZFMF0uitB10BIvzPgLz8XFzK/WtQRwSjYKUHtckm/nUIvPw4EM7/edQ/B4IonQajySr9uHy8/tHA0v+w9DcFReyVBgzpLv7uLLj/peTS/HAYLwbdGI0Fym0y/yhQuPwKlNb8kzgjBVBYhQbDTTb9HRi0/loI2v/jDBsHt4x5BDLhOv3SCLD8rETe/HKYEwbWwHEFhaE+/+D4sP9iiN79yFALB6F4aQZZ8UL/lkCs/hGk4v3hNAMHJSRhBXlBRv1KcKj930Ti/VCf9wBAQFkHUXFK/shMqP9mfOb8oEvnAkIkTQUSaU7/xGyk/1W06v+hS9cAvZRFByyxUvzNWKD/7pzq/UFLxwLIuD0FVE1W/fiYnP/YFO7/gZu3AxLYMQUq6Vb9n5iU/CR07v5GYCcIIeMbBiSolv5tERz9KLhi/deQIwuprx8EasSW/rPhFPwdEGL+DMAjCVmLIwd8MKL9dkkQ/gB4avyODB8JaVMnB6mwqv3+vQz/RKBy/UtYGwlNKysGIjiy/NZxDP7c9Hr99IwbCrk3Lwaf8Lb+qvkM/1rQfvxZwBcJ/VMzBkZkvv9QfRD/YcSG/68UEws9UzcHETDC/26BEP6BSIr9MGwTC6VzOwVUxMb+SMkU/XGsjv3BtA8KfY8/BSSIyv6yGRT8JeyS/8MYCwhJ80MGxozK/ALJFP4wMJb/iKQLCK3zRwf5kMr/T2EU/v9skvziGAcIygdLBsb4xv/JARj+iWiS/Ku4AwryE08E2KDK/6OlFP/ikJL9GUwDCy33UweVmMb/E80U/s+Yjv8iF/8H6jNXBoGMxv36NRT9/viO/RGH+wQN81sGU+DC/r+dEP6IXI7+GQf3Bem/XwYnOML/cQEQ/nbEiv9oe/MHmV9jBI4Mwv3KfQz9yLCK/NBf7wexR2cHYJjC/5e1CP/OQIb/oDfrBRjLawUEtML+xXEI/SWMhv3wG+cFmItvBdQowv5D6QT+dHSG/qAj4wT4K3MEL2C+/ZAVCP4LvIL/qB/fBMuXcwQVEL79xh0I/Kosgv1AS9sFIuN3BkKYuv9XjQj/iDyC/FCb1wciN3sEBwS2/P6JDP5xvH79QQPTBhlvfwT06Lb++L0Q/wRsfv4ZR88GfK+DB6qksv+YKRT9o2R6/Tn3ywYrq4ME7BSy/6L9FPw51Hr98qvHBiq7hwSguK78ApUY/6O4dv5y18MEqdeLBMPwpvza+Rz+gHx2/xuvvwaYe48FZ0Ci/MbJIP7xIHL+QCe/BGtjjwZ86J79b3Ek/4hkbv6Qs7sGog+TBalEmv9nSSj+OhBq/tnXtwboy5cHzZSW/OnxLP3XSGb+wk+zBfO3lwTHzI78mekw//7QYv2jV68GhjObBDKMiv9xCTT+0pxe/bhHrwao658GyTiG/Zy5OP/6gFr8GR+rBgevnwbpyIL/O904/uAYWvxx56cFci+jBzc0ev7unTz/1mhS/WLzowSwm6cFnbx2/RYNQPxCDE79k6ufB6NbpwdTbHL/o7FA/JRETv6gw58E/Z+rB+6Ebv9qjUT9DERK/1mfmweUV68Fz6xq/q0NSP+SMEb9MsuXBVKrrwcrnGb90bFI/L5YQvwTg5MGXSOzB7kQZv5ZiUj+V8A+/zjHkwfTj7MHiMxm/n5dSP/DvD7+OcOPBX3vtwUL5GL8BTVI/dJ4Pv1SV4sFZJu7BUUwZv5M6Uj+a6w+//uLhwdK17sGj9Bm/Q3RRPw1WEL+kKOHB/EDvwWgkGr9DNVE/GXIQv+Jz4MHU7u/B1kwbv1adUD8QahG/PsDfwYqN8MHl3xy/NHRPPwieEr/WCd/B6jfxwUZBHr+snE4/ZbkTv1he3sFvuPHBFnIgv8sMTT/4ZhW/UJ3dwS9q8sGzViK/Fn5LP+7GFr9k+NzBMATzweGNJL/fb0o/l6EYv6Iy3MEoufPBtZMnv/CqSD8FCxu/JojbwQNO9MERQiq/h/tGP/8hHb9C39rB0xn1wa2zLb8QJkU/QOsfv44W2sH9u/XB538wv/vRQz9fOyK/JIrZweWR9sG7WTS/Z0hBP2UlJb+E4djBJk73wamzN784VT8/0cInv+462MGWFfjBP2I7vyYYPT8blSq/tIrXwUT0+MFpaD+/89I6P1G3Lb/Y9tbBWsD5wfEMQ7/PpDg/DH0wv7xM1sE6m/rBZiNHv4n5NT9+fTO/dKXVwVSD+8HVGku/xT8zP/ZSNr8OF9XB+Xb8wR7QTr+JjDA//+I4v+Jm1MF7Xv3BN3RSv6R2LT/iMDu/8NbTwRxi/sG5qVa/PUIqP6b6Pb+EKNPB3mr/wXRWWr/r+yY/lCtAvzqb0sHALwDCxeFdv+biIz/OR0K/GtXRwYa5AMKJ8WC/5gghP0P/Q7/MQ9HBt0cBwqhNZL/T8B0/od5Fv+gJ0cG3wwHCigNnv1e7Gj+FA0e/tDnQwUJjAsJMtWm/Da8XP1oxSL8u0s/BZvkCwqovbL8jtRQ/CytJvzAgz8HigQPCw15uv22pET+ey0m/SrzOwUIiBMKJ43C/dMQOP9jNSr9diQPClyj/wYQgzb+gW2y8kjmCv3yYA8KNdAHCn+HMv6gDmLwsboG/+qsDwtdZA8IAb82/YOW8vNxkgb/xuAPC70MFwqtLzr8wN+W81ZyBvwbUA8LbOAfC2snOv7C1Ar1CloG/HdcDwjw2CcLACM+/PIAavaQOgb8O4gPC0zMLwnZ1z7+Udi+9RMyAv5TrA8KCOg3CgZbPvxwbTL3h+H+/7vEDwt5FD8KpjdC/3NBuveecf78c/gPCKlIRwnNu0b/A2Yi9Twx/v4P8A8KrYBPCiHTSv6bUm708jH6/sQYEwtpwFcKQLtO/LgSsvbnPfb/S+QPCVYwXwtwq079saLO9YMh8vyEEBMKKnRnC7lrTv2pRur0FOHy/sgcEwkmrG8K2MtK/SDOyvUoIe7+ACATCWc0dwtQN0r8AeK697T97v/sMBMJT1x/CRRbRv5qppL3fqHq/Bg0EwpLnIcL/ZdC/asCZvb7Ber9cCgTCNvUjwn0zz79ULY+9R855v9MXBMKWDybClG7Ov9gphL1Lvnm/XSkEwhIQKMLKUM2/DJJxvR4Leb+hIQTC9ycqwiatzL8wml29BRZ5v8wlBMIDMSzCTLbLvwRdR72Voni/9jEEwiw+LsLN/cq/GBMwvW63eL98PwTC/D8wwqodyr+IOx69MyV4v74/BMJFPjLCS0nJv1wICL2k7ne/G1cEwtxANML1uMi/cEH2vJyld7/0WgTCOTc2wrpryL8IUdm8uPZ3vxxkBMJhLTjCwuzHvyiIwbzRvXe/oHYEwl8eOsIPqse/KNapvFn4d7/FegTCjxM8wmQEx7/QLIu8mKl3v7WJBML++D3CZIPGv7CCWLztone/HJ8EwhfjP8LXLsa/sBctvFWpd79wmwTC5r9BwmP1xb9A4/q7RfR3v/zMBMKImEPCJuvFv2BR3LsXHHi/absEwiR8RcLq6MW/ACizuxRoeL8X3QTCj0lHwoRvxb+A8oS7rNZ3vzvvBMKcGknCy5nFv8CRl7tcBHi/IxAFwgjzSsIX28W/8Ie8u/U6eL8WFgXCHrdMwo7uxb/ofgm8/bd3v0k7BcL0gk7CxhHGv+jKL7yNZne/kTkFwvhCUMJ6nsa/2NlvvMR7d7+KVgXC8v1RwlAIx79kvJq8pzZ3vw1oBcLxxVPC2pXHv9iExbxL9na/e4YFwlpxVcIjHMi/qssBvQgNdr9xggXC2i1XwharyL9+Lx69lFt1v5+dBcLZ2VjC/SzJv7L9NL3c53S/PcAFwmuHWsJMR8m/1JNKvUDAc78owgXCckZcwiWoyb8AaF69Nzpzv5zSBcI74l3CeuXJv9Iec71GYXK/DOYFwr51X8Ksl8m/Jrh6vQhUcb/U8QXCKiphwvpByb/sm369WHRwv/YNBsJ3xGLCennJv00shb1KIHC/px4GwrZVZMI72si/h/yBvRRcb78cMAbC39llwv7cyL+bUYO9rDZvv1YxBsJJc2fCFK3Hv4akeL3R3m2/RkEGwi0NacLT6sa/xOZhva3cbb+lRgbCLJVqwhMVxr/ca0u9FLBtv6k+BsJ8J2zCOFLFv4D6L72h8W2/ZEwGwm6ybcIJXcS/UAATvTrqbb+XNgbCljlvwlL+wr/ss9K8yNxtv41MBsJnwXDCVHHCv2TembzEhm6/iloGwixOcsKwKcG/8Ozau+IPb784YQbC3tVzwvYpwL/gn107Q5pvvzZuBsK/bHXCXUO/vzgJdzz0sHC/u3EGwl/pdsLQ+72/YAjgPG0rcb97awbC4HN4wrxUvb9Eqh09g4pyvxRoBsKECHrCmim8v0C4Uj0IU3O/f2wGwrKIe8ItMLu/EUGBPQcmdL/bSAbCFSp9woQ0ur/4zJk9DgB1v3lzBsI8tX7C/OK5v+ayqj1FQ3a/zVYGwosxgML/cbm/u1O9PSl4d78DYgbCWfyAwviiuL+eqtM9k1p4vwRVBsKowIHC1QW4vw1z5T2+GXm/aVIGwjmJgsKoibe/gEzzPaSoeb9JbQbCQleDwswOt78uMwE+xVp6v+VRBsJ3L4TCitW2v380Bj5t/3q/LGIGwhn0hMJHZ7a/FvAKPv0qe7/OVgbCk8GFwtpYtr8ong4+xNl7vylkBsJXkobC2Vy2v+oAET6UZXy/P7YWwmlLzsFs3hdAQ7w5v+GmMj+z1BbCwgXOwUi4F0DQZjq/ToUzPzz1FsLxws3BRLgXQOHoOr8mujM/nBMXwmCCzcFEkRdAkGg7v2GLND+ULRfCBETNwdl1F0ACkDu/Pwo1PyBFF8LCFM3BGYUXQN4RPL+JATU/q14XwrbozMF8tRdABOg8v9SUND/EdBfCpLXMwa/iF0B8fz2/Lhs0P1qOF8LajMzBgkUYQKMyPr9J0zI/7qgXwlJrzMHg1xhAkVw/vwf6MD+rwhfC00/MwRtxGUAAOkC/NuYuP+7XF8JSKczBIBMaQNaPQb8q3Sw/+O0XwsINzMGnrxpAVPRBv1+NKj+uBBjCq/XLwcR8G0Dc+0K/tLgnP7UkGMIB2cvBMxocQDzZQ7+JkyU/nDYYwh/Iy8FkwhxAE1JEv+0eIz/2RRjCxq7LwbhQHUCUg0S/DfkgP39aGMJGksvBPdAdQMiXRL/9BB8/5HAYwid/y8GXMB5ArJVEv+iFHT81gxjCcm/LwTt7HkA1vUS/6mscP3CRGMLyX8vBKbEeQEnFRL9ImRs/KZ0Ywo5Wy8Hw3B5A5fdEv4v9Gj+mpxjCs0/LwbzwHkA6I0W/Gb4aP6G2GMIyRcvBqvoeQLmLRb+Buho/IL8Ywvw/y8GZEB9ADwhGvxOOGj+uwRjCNTfLwc7kHkDrFka/lUAbPzvNGMLqMsvBasIeQFhFRr/S2Bs/JdgYwuU6y8Hjjh5AfjRGv4OfHD8d4BjCEivLwXdTHkAa90W/TXYdP7nfGML6K8vBlhUeQBKzRb+VVB4/O+EYwiAky8HUxR1AwPpEvyZRHz975RjC0hnLwQWAHUAogES/jTsgP0fsGMLNHMvBfiYdQC26Q79eWSE/bucYwgQUy8EQ2hxAGSlDv5FVIj856BjCjAvLwY2HHECDZEK/DFcjPwnqGMKe/8rB+EYcQPPWQb+2JCQ/L+wYwhT8ysFD5xtARj9Bv2hqJT/J9xjC8urKwc2vG0Az3kC/pyMmP4bxGMIe4srBQGsbQAKgQL8IHic/cfAYwibWysH+PhtAHVlAvzy0Jz+h8hjCTsfKweIMG0AaOUC/ZnAoPwvyGMJ4tMrBCNUaQCQlQL80SCk/gvMYwmq2ysF8qhpACaI/v8fAKT8Z8hjCmKnKwfqNGkAntj+/gDoqP6TyGMIqlsrBc1oaQKx9P79X8yo/zPIYwvGNysHxUhpAHko/v7P9Kj+6/hjCXIPKwRg2GkAtRD+/AG8rP6L8GMLYe8rBcSUaQPArP790qCs/EvoYwvdzysFjGRpAfhU/vyXQKz9MARnCHmfKwX4GGkDuPD+/GSssP9r9GMIxZMrBCf0ZQCIVP7+1QSw/jv8YwplRysFX+xlAHAs/v6VELD+BAhnC7lLKwWEDGkACEz+/cicsP1gIGcI3RcrBvgUaQIE4P79mLCw/oAoZwvxUysGqLxpAnZQ/v6KnKz/pFxnCqUbKwcIkGkA52j+/GO4rPz4VGcK+PcrBfUoaQA/6P7/gYis/5B4ZwrEqysE7VxpAchpAvyM8Kz+BJxnCLCbKwWGRGkDxwUC/rJIqP5c0GcJzHMrB7bcaQC5JQb9KKyo/8jgZwuQxysG+4xpAC89BvwWuKT8zQhnClirKwREZG0AvU0K/sgkpP5dMGcLnKsrB81sbQCkfQ7+ISSg/XlwZwoEjysGLlRtA4ANEv1C3Jz+AbxnC3B3KwfXhG0BF/US/d+AmP8h+GcIuDcrB1C8cQJHIRb8X8iU/8YkZwq0UysEHmRxAy7NGv+6gJD+2mxnCwhjKwf31HECU3Ee/TZYjP7meGcIUHsrBYVgdQHCBSL/LRSI/lrcZwlgVysE5rx1AcnZJv5w/IT+XvhnCDP7Jwd0PHkCL50m/NuMfP3zHGcJ9GMrBjX0eQP2fSr8xax4/SdcZwl8TysGk1x5Ase5Kvy0dHT855hnCqQPKwassH0DX7Uq/m8gbP1n5GcLeCsrBxoofQNjvSr9dURo/KQoawnIXysHA7R9Adf5Kv5/LGD+YGxrCPxjKwRo2IEAcoEq/eIwXP4wxGsIvNMrBwZYgQPtcSr/C9hU/vkQawk86ysEx5SBANvxJv6WgFD8sWBrCD0jKwZY+IUCepUm/lCMTPx8qGMJ0msvB6xgYQFM9Or8T8DE/o0UYwgxay8Fv9RdAguc6v0jDMj9QYxjCYhzLwVj2F0Dhazu/CfUyP+5+GMLV4MrBPdAXQI3vO7/uwzM/DZYYwlenysFVtBdAgh08v0pHND/iqhjCw3zKwQLCF0BspDy/uUY0P83BGMI4VcrBRvAXQCB+Pb+b4zM/I9UYwmUmysHwGhhAthU+v/pzMz8S7BjC0gHKwdx6GED4xz6/IzcyP/gDGcJX5MnBZAoZQNDtP78tZzA/FBsZwm7NycF+oBlA3MRAvxJdLj/SLRnC5arJwXQ/GkA+E0K/31wsP0RBGcI0k8nB59kaQFhvQr8NEio/bFUZwsp+ycEhpBtAMW5Dv0JFJz+9chnCymXJwU8/HEAnQkS/UyUlP0eCGcKDWMnBiOUcQEG0RL/xtSI/SI8ZwotCycFLch1Abd5Ev5iTID9UoRnCIinJwSbwHUDw7ES/G6QeP6C1GcJLGcnBPU8eQLTmRL+SKB0/wsUZwmIMycENmB5AIwpFv1cUHD/B0RnC1//IwebMHkBsDkW/rUQbP8HbGcIK+cjBUfceQNA/Rb/VrRo/KuQZwr70yMEKCh9AzGlFvyFyGj9m8RnCtOzIwcQSH0DM0UW/CXMaPwX4GcLT6cjBVycfQG1ORr/8Sxo/MfkZwk7jyMEo+x5ABGFGv08BGz/bAhrCIuHIwfrXHkADkka/j50bPzIMGsIk68jB6qQeQMWGRr9hZBw/mRIawsbdyMG1aR5ATU9Gv3Y8HT/bEBrC1uDIwZssHkAgEUa/0hkePzYRGsIZ28jBYd4dQM5gRb9aEx8/9BMawt7SyMFimh1Aau5Ev6X5Hz+HGRrC/NfIwXBCHUDlMES/fhQhP4kTGsJ20cjBR/gcQGmpQ79ZCyI/IhMawgvLyMEVqBxAC+xCv5IGIz/OExrChMHIwT1qHEAvaEK/IM0jP7oUGsIdwMjBLA0cQFTYQb+HCyU/PR8awkixyMG52BtApIBBv3W8JT8zGBrC2KrIweCWG0BhSUG/1q4mPwkWGsJOocjBrW0bQIsJQb+kOyc/YBcawsGUyMGCPhtA9PBAv/fuJz/eFRrCXYTIwakJG0DP5EC/070oPy4WGsJtiMjB4eAaQJpmQL+DMSk//BMawjR+yMFMxxpA9YBAv+KhKT+eExrCjm3IweiVGkCKTEC/5VMqP+USGsL9Z8jBI5EaQK0eQL+UVSo/yx0awqZfyMFcdhpAYRtAv6W/Kj/CGhrCr1rIwa9nGkCfBkC/lPIqPxMXGsJ0VcjBjl0aQN/vP7+AEis/Ux0awuFKyMHSSxpA8hhAv1hpKz/QGBrCUErIwepDGkC+7z+/TXkrP7AZGsIROsjBlkQaQOrkP793cis//Boawng9yMFDTRpAaeo/v8VRKz/cHxrCKDLIwShRGkBtDkC/6E8rPy8hGsIlRMjBf3waQPhnQL84xCo/Ji0awr03yMGIchpAWKtAv+AFKz9HKRrCpjDIwUuZGkA7xkC/inQqP7YxGsKAH8jBzKcaQFLmQL+KRio/wzgawsEcyMHw4hpA34pBv4mXKT+nRBrCrxTIwf8KG0C+EkK/9SkpP4FHGsLWK8jBMTgbQFSWQr/xpSg/l08awiUmyMHrbhtAXRxDv1r8Jz9dWBrCMSjIwRazG0AT50O/ATYnP59mGsJsIsjBcO0bQM7KRL/AnyY/THgawrEdyMHYOhxAJMNFv9bDJT/5hRrCXg7IwQmKHECXj0a/DdAkP72PGsIUF8jBQvQcQNN4R79zeSM/rp8awlIcyMGHUR1AL6BIv0BsIj9YoRrCuSLIwcW0HUBOREm/qRchPyS4GsLvGsjBOwweQNo1Sr8mDSA/0L0awicFyMHwbR5AmKVKv8WrHj/exBrCTCDIwbnbHkAbWku/lDEdPwjTGsIWHMjBEjYfQHKnS7/14Rs/hOAawngNyMEzjB9AKKNLv/WHGj8A8hrCwBXIwZ3qH0AGoEu/7A0ZPy4BG8IwI8jBkU0gQFesS7+lhxc/FhEbwpokyME7liBAzUpLv6NGFj9mJRvC+kDIweH2IEDZA0u/HLAUPxw3G8KER8jBM0UhQLShSr95WhM/4EgbwulWyMEFnyFAs0pKvwzcET+KXx/CfvG+wUbIH0BnFE6/7mgaP/t3H8Izvb7BebYfQCzSTr/q7xo/hJMfwraKvsHPuh9AbVdPvxALGz/frB/CQFm+wWSTH0BSuU+/rcobPz/BH8I3Kb7Be2sfQH6yT79YaRw/JNMfwo4IvsHxZB9AXvJPv02ZHD8k5x/Cu+m9wbN1H0DdcVC/nIAcP6/2H8Kwwr3Bm38fQLaTUL/9Yxw/zAkgwkKmvcHGtR9Akc9Qv0SdGz9MHSDCHJC9wXkbIEBLbVG/QjcaP0YwIML+f73Bx4sgQIu7Ub81jBg/tD4gwsBlvcHFByFAkIdSvz3bFj+qTSDCwFS9wRmFIUD4dVK/st0UP2tcIMIuRr3BjjAiQB0WU7+NYBI/MXUgwo4zvcFMtSJA95ZTvxN1ED+sgCDCaCy9wZFFI0Af0lO/hkYOP9KIIMINHL3BU8QjQLvTU791TQw/tZUgwloGvcHtMyRAMs9Tv/+PCj9epSDCvPy8weiIJECwtFO/pDYJP7GxIMLa87zBv8ckQETRU7+6RQg/MLkgwtXrvMEA9yRA/tFTv6iKBz8avyDC1Oi8wfwaJUBGCVS/GAwHP6HDIMKw6LzBxiglQPIwVL/q4AY//swgwtjkvMH0LCVA6aBUv5vwBj8xzyDC3+S8wbQ5JUAWHVW/weEGP4vNIMKa4bzBgA4lQPI/Vb9Tlwc/rtIgwt7hvMFR6CRAzYdVv+VDCD9C2CDCVe68wQy6JEDwoVW/nQMJP0LbIML45LzBgIIkQNiTVb+93Ak/B9UgwjTsvMEMTiRAsIVVv62pCj++0yDC1+i8wYoKJEBcFFW/V5ULP3LTIMLB5LzBdtMjQEnqVL+QZAw/I9YgwpfsvMHXiSNAcXtUvzJpDT/nzSDCEuq8wWRMI0AeO1S/L0sOP57KIMKO57zByAojQMK/U7/BKw8/yMggwrTivMEL3SJA7YJTv+vPDz+bxyDC1uW8wbiRIkD3MlO/fOQQP4XPIMIi2rzBdGkiQNMYU79/fRE/WMggwkXZvMGINCJA8wlTv85MEj/5wyDCYNS8wXQVIkAY9FK/gcISP6rDIMIDzbzBZ/EhQJj6Ur8cVRM/EsEgwgDBvMFMxCFAQgVTv4ANFD9zvyDCFcm8wXejIUDYlFK/nW0UP+a7IMJewrzBQ48hQLO7Ur8byxQ/FrsgwhO3vMHMYyFAt4hSv3NpFT+/uCDCk7i8wTVkIUCOXVK//FkVPyPCIMJFs7zBkUohQIxWUr/FvhU/670gwvCyvMGqOyFAjDpSv6rxFT99uSDC8rG8wQUyIUBfEFK/2AoWP+y9IMIPrLzBwh4hQCYuUr/gYRY/BrkgwomuvME3FyFAkOxRvwZrFj9GuCDCqaK8wXMaIUDwzVG/JFQWP1S2IMLdp7zBcR0hQCexUb/VPhY/WrogwrufvMFIISFA8rpRv5AyFj/UuiDCiLK8wUtGIUAr71G/vK4VP/PEIMJuqrzB0zghQMsJUr9l7RU/0b4gwtSkvMGkXCFAyfhRvxRYFT/qxCDCbJW8wWhoIUAe7FG/zCQVP63KIMKylLzBWp0hQFtkUr+3dhQ/+tIgwoyNvMHWwCFAcsNSv4gGFD/X0yDClKW8wV/rIUCEGlO/YXcTP2TZIMKXorzBmBkiQLpzU7/+2RI/tN8gwiqlvMHOVyJALgpUv3oPEj/J6yDCIKC8waeLIkBOwVS/X3gRP8T6IMJNm7zBDc4iQFKBVb8gqRA/nAUhwvuMvMEJGCNAQyBWv8OwDz9TCyHCl5a8wd54I0C5y1a/GWAOP90XIcI8nbzBe8YjQBC3V79Nbw0/OhYhwgSkvMEKIiRA7ChYv6UhDD9sKSHCzJq8wexsJEA831i/9ioLP5gsIcJ+ibzBrMgkQK4UWb+Bygk/mjAhwqCivMF5KCVAC4tZv9JsCD9+OyHCCZ+8wXJ5JUA8tVm/qzQHPzZGIcJBk7zBmsolQBh6Wb9f3wU/CVUhwuCavMHwICZAeUpZv2d5BD+GXyHCRam8wUZ3JkALO1m/Ix0DPwRvIcIMq7zBsbsmQFnAWL806wE/vX4hwkDHvMG8FidAGV5YvyBnAD95jiHCPsu8wZBdJ0AO81e/gGP+PsGcIcLY3rzBlbMnQP+TV7/biPs+7n1AwudvRsGuRzq/aNQ6P9yeKL+qoT/CoQ5JwSz8OL9AXTo/2yknv5i9PsJVq0vBdD85vyQ+Oj+GYCe/qOI9wn46TsFDXTm/1PA6P6PBJ7+5Cj3CstNQwc6LOb/EfTw/LoYov6QsPMIzhFPBQZQ5v0DyPT/SGym/LE07whxHVsF/Yjq/CU0/P4ptKr/ieTrC0ftYwZVcOr8wVEA/4Msqv6agOcIPw1vBixc7v4/2QD+SxSu/Lck4wp+LXsEo2Tu/YA1BPxORLL/S9TfC9GRhwWNiPL9HrUA/OPYsvyMuN8K3HmTBI6Q8v9oCQD+e9iy/Q182wgPiZsHZCDy/wLs/PzM/LL8vnTXCSKRpwafmPL8fiD4/Eqcsvw7YNMI0R2zB2Zo8v1DZPT+eFyy/CRs0wmgfb8Fp/jy/LrI8PxYJLL+5YTPC1bRxwbssPb+6UTs/2q4rv+qkMsJ1U3TBY049v6Q2Oj+nYiu/nuwxwqrddsHabT2/VA85P2IPK7+ePzHCl4x5wahaPb/VFTg/o5sqv9yPMMLi/nvBrGE9v6WONz8obiq/deMvwiiZfsEQfz2/7ks3P01xKr+AOi/CUJOAwc9UPb/+jTc/NGEqv36RLsJSyYHB6PY8v2teOD8eVSq/F+otwlz3gsHrSTy/UBI5P3jvKb86TS3CBieEwWOLO79NPTo/xKUpvxyzLMI9TIXB30g7vwsIOz+HsSm/ZxUswqNxhsEN9zq/zSI8P3fMKb+AhivCm4qHwVqPOr/+/jw/YLkpv6j5KsJjpIjBDw46v1DvPT8YlCm/cFwqwoO8icEqATm/UAI/PyrwKL9x0CnCB7mKwQIUOL//BEA/7mQov+A7KcJRxIvB08k2v0gTQT9RgCe/sbAowhm7jMHCADa/qv9BP0wPJ79ZOijCaLaNwbhaNb8ZmkI/cKImv5CoJ8KWvo7B4e0zv0qrQz/6mSW/licnwiKtj8FQtTK/DHdEP5urJL8uqibCJJ+QwUVOMb9PlEU/k6sjvywrJsJslpHBXS0wv8KMRj+S4yK/PKklwkp7ksH6ai6/pExHP/JkIb+qLiXC3VOTwaiDLL8rckg/muQfv2mrJMJ1RpTBhYkrv+A6ST9JMB+/EjgkwlQSlcEd2Sm/cSdKP1DRHb/VuSPClu+Vwdl7KL8370o/PLgcv7RJI8JCv5bBP8Ymv2JHSz+GIBu/0MciwoGOl8ETOSW/hG9LP26hGb+bVyLC+FWYwZFrJL9vm0s/HeMYv3PgIcIDFZnBPzYjvwZeSz+ymhe/qmAhwn7tmcH2xyK/rRdLP9oVF78O8iDCwpyawadGIr9WQEo/wk4Wv9eBIMIiR5vBwmEhv+XdST/hSxW//w8gwmARnMGkQCG/QBxJPwjsFL/7nx/C8cycwbfBIb9Qy0c/Z/0Uv38wH8JWiJ3B09Yhv8QQRz8x1RS/hcgewiIbnsHqsyK/BHxFP7QpFb/6VB7CftuewTlII79kRUQ/xFQVv63vHcKyiZ/BMhkkv2+HQz+J4hW/I3YdwiBBoMGjeSW/wGJCP+HZFr9WCh3CEOCgwYCVJr+/nEE/AK0XvzmdHMJDraHBmo0ovyi6QD9JTRm/liIcwr1OosHykym/X7lAP/RNGr9hzBvCBxmjwVqeK78Uqj8/JPAbv7pnG8K/z6PBhQotv8l5Pz+PRB2/vPkawreKpMHMLS+/2iE/P2w/H7/djRrCrVylwTJbMb+d+j4/nlYhv1stGsLxIqbBpy8zv80cPz/xMSO/IcgZwk/vpsGThDW/N9s+P19pJb9uXBnCAcGnwcy0N7+cuD4/OIknv4v/GMKqk6jBYIE5v6itPj8MUCm/SIcYwgVrqcEEZzu/NEA+P1YLK7+PLhjCJV2qwfXXPb9gsj0/DkYtv7nIF8I7VqvBle4/v6/xPD/TEi+/o2QXwtEtrMFr9kG/Z4U8PzzyML/S5hbCcimtwYV7Q7+XLTw/ylYyv6OJFsLVIq7BmG5Fv7ZMOz8J8zO/MlgWwoMIr8GFLke/2DU6PxVFNb863hXCmSqwwVPUSL8z7Dg/zWY2v0qOFcKnPLHBL3FKv4WjNz+Ofje/kiMVwpwpssG+yUu/3iE2Py84OL8a4hTCaFqzwU+UTb/PjTQ/F1s5vz9sSMIXZinBkrE6v7I2Oj+Pyyi/YXZHwshULMHdODm/svE5Pxc9J7+Id0bC7EMvwfNfOb/8Cjo/Pm0nvxSBRcJnJjLB3V85v7jqOj/mwSe/W49EwgAXNcF8hzm/EZE8PzWJKL/ClUPCICA4wQKGOb9uDT4/7hcpv5GbQsKMPjvBH146v2lqPz9adCq/OK1Bwh1RPsFoRzq/vGtAP5u/Kr88uEDC2nhBwZT7Or/2BUE/Wa8rvxHFP8KAoETBEZM7v+oSQT+qTCy/Z9U+wnrYR8EyBTy/6JxAPyWSLL9Y8T3CyPJKwV48PL9W0z8/A3wsv/8FPcIeFk7BAnE7v6B2Pz9VjCu/Byc8wkw5UcGDJDy/MyM+P7S9K7/IRDvCkjpUwQi4O78GZj0/jwgrv/lpOsLCclfBJeY7v7EyPD91wCq/AJM5wkpnWsEB7Tu/DMs6PwQ9Kr/9tjjCpGVdwV/KO7/wwzk/kLUpv9rgN8KsTmDB/rw7v8KsOD8aPSm/ahU3wjtdY8E8gju/rck3P/yrKL+ARjbCqipmwbJUO7+4aTc/W1oov5x7NcKGI2nBjkQ7v/NWNz9JQyi/SrI0wuANbMFS7jq/nMk3P0gaKL+c6DPCVthuwbRwOr/p0Dg/BAMov9ghM8JFknHBvqg5v/mzOT8blCe/AmQywgFMdMFM0zi/FBY7P0dHJ7+fqjHCyu12wYuIOL+JCTw/I1knv9XsMMJOjXnBnDQ4vyJIPT/dfSe/Oj4wwmAafMHwwze/t0g+P09uJ79qkS/CrKB+wW9EN7/iRT8/RU4nvxrWLsKmkYDBAT42vwBmQD9AtCa/qiouwtG3gcFHWzW/eG5BPzE0Jr/rdi3CU+qCwa4mNL/nfUI/LmQlv63NLMJBB4TBvVYzvwBtQz8I7CS/oTkswi8phcHdwDK/1wJEP9yMJL/tiSvCO1eGwXxdMb8NE0U/J4wjvyzrKsJYaofBoSowv3jiRT/QoyK/OFAqwrN/iMFjvi6/NwNHP2CeIb/MtSnCOpqJwRmOLb/E+0c/ucUgvyIXKcIyoIrBxMIrv1q3SD/COx+/GH8owrGYi8F9wSm/y+RJP7iiHb854SfC1qqMwcW2KL+5vEo/8+Ecv6ZTJ8J7lY3BrPMmv62kSz95bRu/mLomwkONjsF0dSW/x2NMP3svGr8nMSbCR3qPwRyiI78NtUw/kHcYv0GVJcKvZJDBqOQhvxPnTD+zyxa/6AslwvREkcE0+SC/Of5MP7noFb/ZfCTClRqSwcmeH79UtUw/o3gUvzznI8IJDZPBVxYfv1lWTD+o0hO/518jwovRk8HPZB6//nJLP/zZEr9z2iLCy4+UwSNUHb869Uo/k6QRvxhSIsIMcJXB5gMdv5AUSj9HDhG/Fsshwsk/lsGhYh2/NbFIP776EL8jSSHCRAqXwZheHb/+2Ec/JrIQv8TLIMJYrZfBKQAev4oxRj+QyRC/+0Qgwud+mMGUcx6/pPREPw/VEL8OzB/Cyz+ZwTolH7/NJkQ/PUARv6o+H8KUBZrByFAgv4sIQz+dBxK/JL8ewnW0msG2SSG/S1NCP/m+Er9YQR7CkI6bwSMjI7+Mg0E/m0cUv0iyHcL2PZzBFv8jvyimQT9SKRW/cksdwlUTncET2SW/n8tAP5SuFr8M2BzCl9SdwXwcJ7+hz0A/pOsXv2xZHMJYnJ7B7i0pv1yxQD9A5xm/HN0bwsh6n8FMOCu/ZtVAPx30G7/ibBvCFEugwSnoLL+3T0E/2Mcdv3H5GsI9IqHBCSovv0VbQT+ABiC/an8awtz/ocF6QzG/wJhBP/cwIr/GExrCzteiwTbpMr+D8UE/U/Qjv+qLGcLyvKPBxLc0v6LrQT84vyW/CigZwre4pMFGEDe/OrtBP3YFKL9zthjCO7+lwUEOOb8iWkE//N8pvzJEGMIXoKbBA/86v8FZQT/X0iu/2bcXwmClp8HnaTy/n2NBPxZELb9xURfCHKiowZ0+Pr+90kA/1uQuv6gUF8L6mKnBZf4/vyAVQD/wXjC/G48Wws7FqsH+m0G/ZBE/P1KaMb+KNRbCYN+rwVgqQ7+CBz4/+cIyv77BFcKo1qzB8X5Ev2C+PD/NljO/aXkVwkgSrsF5P0a/v1g7P7zKNL/mEDrCo8yzwZTYGEAXdj2/vTYwP6MbOsLTpLPBkucYQIYIPr9ENDA/2yk6wgJ/s8GoExlAt3Y+vx6uLz/SNjrCuFqzwRkZGUBd+D6/MssvP1w/OsKqNrPBNBkZQJQ+P79k5i8/E0Y6wiAgs8HENhlA29I/v1CpLz+sTzrChguzwRhpGUDKr0C/jDQvPwlVOsLX7bLByYoZQO4yQb+I3y4/ql46woDassGa2hlABdFBv2TaLT+vaDrCP8yywWFVGkCQzkK/GUwsP3lyOsKgxbLBRtEaQOl7Q787mio/png6whqyssGwUxtAuJlEvwj4KD+SfjrCuqiywa3SG0BAzUS/LQwnP0qFOsLJobLBOHscQLOYRb/ysSQ/ypU6wn6WssE6/BxA3ThGv7PmIj8KmTrCOpaywUeMHUChlUa/fMcgPwyaOsKSjLLB3gUeQHGkRr9z5x4/sJ86wkB9ssFOdR5AmqJGvy0rHT/DqTrCa3mywdnOHkDEkka/CsIbP8ivOsKcdbLBDhEfQHirRr/kwxo/SrE6wu1yssEmSR9Ajq1Gv2PmGT/NtDrCanSywWR5H0BC5Ua/QDoZP1C0OsJLebLB05UfQOoUR7/F2Rg/d7k6wjp5ssEjrB9AT4pHv96oGD/TuDrCVHuywW/OH0CEFEi/Sk8YP/S0OsKofLLBtrofQFlHSL9trhg/yrY6wtWAssFjrB9AgZtIv2wDGT8mujrCko+ywUuXH0ANyUi/b2YZP5+6OsL/irLBhXUfQG3NSL8R7hk/6LQ6wpCSssEvVB9AgtNIv6d0Gj+SsjrCT5OywQkmH0D2cki/eAsbPyOvOsIUkLLBSgIfQE5eSL+rkhs/4bI6wjSassHExh5AFgBIv2FfHD/TqDrCrJqywcWYHkCh5Ee/Iw0dP6qmOsLumbLBg2MeQKd7R78bvR0/HKU6wpmWssGfPh5A72pHv19KHj++ojrCCpyywb36HUAPMUe/A0UfP8KqOsI7k7LBZd0dQHhIR79kwh8/2aM6wkCSssGirB1AdGZHv9+PID/JnzrCCpCywe2SHUA1fUe/xP4gP8OfOsKEiLLBV24dQDavR78GoyE/Tp06wg2AssE1RR1A2+1Hv0ReIj8tmzrCG4uywd8fHUAVpUe/LNoiP9SXOsIghrLBhwsdQIP1R7+mSCM/G5c6wut+ssG53BxAfu9Hv7ICJD8plTrCNIOywVjcHECy7ke/7gMkPxudOsKMf7LByr4cQHb8R7/XfyQ/SpY6wvqCssHrrRxA5vtHv5DDJD8ikDrCyoWywYSkHEAa6ke/AeMkP06VOsKogLLB1owcQMAVSL85UiU/hY06wkOHssEYghxA59pHvy1oJT+wjDrCEXyywZ2DHECYwke/PVklP2yIOsKrhLLBj4IcQLWiR7/pUSU/fYo6wvyAssEBhxxAiKtHvzZDJT+QiDrCFJiywSywHEBy3Ue/fa8kP8+POsLij7LBn6QcQAP3R7895yQ/wYg6wkWLssGlyRxAsdpHv/ZHJD/vjDrCkICywQjcHECw3Ue/FP8jP86POsImgLLBQhYdQFBPSL+nPSM/fJc6wtR5ssHLQR1AZL1Iv+S1Ij8gkzrCmpSywR1wHUB+C0m/bxciP1GYOsI+krLBTqgdQJKBSb9VXyE/N5s6whSXssHG7B1ABRJKv/x+ID9YpDrC1pKywS4nHkACzUq/l9UfP5ywOsKqjbLBuXQeQN2BS7+y3B4/g7g6wuR/ssFKxR5A8SxMv9vTHT9xvjrCDoeywTgtH0DO00y/WGscPwDIOsJyirLBLoEfQHuzTb/TZRs/vsU6wuKOssGS4h9ARBpOv2EBGj8k1jrCOYaywWg2IEBqt06/UeUYP7fYOsK3crLBlJggQK7kTr/eahc/Pts6wriGssGd+yBAbjxPvzX7FT8O4zrCzICywb1NIUBYSU+/QLcUP7LvOsJQcbLBq6IhQIT9Tr9UTBM/lvs6wqN2ssHZ+CFAr61Ovy3cET9lBjvCGoGywRNTIkAfi06/JmsQPyYVO8KIf7LBaJciQGf+Tb+1MA8/uSU7wg+YssHU7SJA2oJNvyq1DT+XNTvCTpeywWc0I0AeHU2/538MP0RDO8IFqLLBZIQjQJ+4TL/fJgs/7FAVwqghJkK5VEe/fI8xPyvfMb/bmhTCkHMlQidhRr/xzjI/v3AxvzLcE8ImwSRC359Gvy4wMz+h1jG/CCYTwpwRJEKHbEa/vX0zP17DMb9VdhLC9mMjQr8wRr9YSzQ/4Nsxv1DDEcKUsiJCT8lFv8npND+PtTG/zw0Rwu8AIkIorkW/j+c1P7gBMr8aYBDCilUhQqjQRL8tGDc/WaAxv9GxD8JQqSBCR3JEv+lEOD+YuzG/XgEPwiIAIELXB0S/alc5P76/Mb+uUw7CTGUfQrrIQr/7PDo/K9wwvyurDcLswR5CRuxBv9vdOj95PzC/Sv0Mwq4eHkKTXkC/LHc7P/LtLr8KWQzCgoAdQpJgP7+/azs/bOstv8OrC8K66hxCtgY+v1CnOz9ZqSy/GQsLwiZNHEJp8Tu/jMM7P/CgKr+UYgrCbbobQnDtOb8WHzw/LMMov4S3CcJNLBtCmto3vwLEPD/w8ia/MhAJwuCbGkLakjW/mrg9P+0LJb94bAjCPRMaQt1ZM7/ksT4/qTQjv+3BB8JZihlCHUQxv/fYPz8OkCG/HCAHwkYGGULPZi+/bxZBPxAqIL9zfwbC8HwYQl2QLb9Qa0I/utEevwTXBcIO/BdC5rgrv4vcQz+wgB2/szoFwox+F0I8nSm/lytFP43eG7+yoATCN/gWQuzWJ7/ZekY/ao8av079A8IMghZCaV0mv2esRz/bgBm/8FoDwg8BFkLnHCW/5bVIPzecGL83zALCtooVQkcYJL/FiUk/R+AXv5gvAsLHDhVCcwUjv/gTSj9d/Ra/1JMBwueYFEKX6CG/bKZKP/ASFr8SCAHC7yIUQpZRIb8KBUs/FJwVvxJrAMKYrRNCWZwgvx81Sz8D+BS/yLn/wWY7E0J/HiC/83NLP52PFL+Op/7BwskSQoxSIL/eR0s/8LQUv5qP/cGPSRJCPQggv7xNSz80bRS/+n38wYTbEUIryh+/WYtLP5hDFL/ge/vBu2wRQmdaH79Ev0s/kuUTvyBn+sF29xBChz4fv5cgTD886RO/VGb5wVmKEELwER+/cVxMPz/QE7+4SPjBex0QQsxSHr/PDE0/GUsTv/ha98G+sw9Cs4Qev9FOTT++kRO/RDf2wZRKD0LjFx6/2P9NPyteE7/YS/XBrN4OQvcHHr/1Ek4/b1QTvyxW9ME0cg5CQZUdvzxITj9a8xK/3lXzwTYHDkL0hh2/ZVJOP13oEr8sZ/LBd6ENQsXRHb+EHE4/miETv6Ry8cGGOQ1CKsgdv8vBTT8t+xK/Tm7wwVzXDEJL2x2/altNP5LtEr/uZe/By3QMQrj6Hb/Lzkw/8d8Sv+iN7sELHAxC1bwdv6ExTD+VcBK/Xp/twdy3C0KnEh6/omVLP3yEEr+MhezB7VULQts/Hr9GjEo/xWsSv3S368F3+gpCGoQevxStST+8ZxK/2L7qwWCpCkKxix6/uqVIP/UaEr/+0enBhU4KQl+uHr9y5Uc/ev8Rv87g6MH28glCIY0ev4JsRz8muBG/lPXnwa+UCUIxuR6/nPJGP1C8Eb848ObB1U4JQoZCHr9i0UY/eD0Rv94K5sEw7QhCriAev3oSRz8dMRG/lADlwWyiCELAgB2/fp1HP7nAEL+qHeTBz0IIQmlPHb+zx0c/yJ0Qv6w148Gc6AdCibscv/9oSD/sPxC/ElPiweShB0IrDRy/zCpJPwDSD78MU+HBaUgHQt15G7/O/kk/8IMPv8xS4MEr8wZCDMsav05SSz8qQg+/KmTfwe6wBkJhAhq/Om1MP37UDr8CYd7BY1EGQvDAGb9rVk0/M9wOvyR03cEl+wVCJKEYvwAkTj9g/w2/Gm3cwaKxBUKk7he/hKdOP1V3Db/4dNvBe1sFQu9aF7+IeE8/FiUNv+qR2sEwCwVCYIoWv44MUD+1gwy/+JjZwQnDBEJz5hW/X25QPw//C79ejtjBN20EQuVhFb//k1A/PIcLv3qs18HvIQRCJiUVv36/UD8xWAu/7rvWwXXdA0LRwBS/+qdQP+ztCr8m2NXBTX8DQg/SFL8ULFA/6NkKvwDs1MFcNANCX4YUv/91Tz/QWAq/yPnTwVrvAkLrlRS/np9OPx4oCr9SEtPBDJoCQrd6FL8q2U0/INIJv2a0KcLfBxxCGq0bQFFrPb+P5yQ/cGArwst0HUJs/RpAVh8+v9HjJz+g9CzCCd4eQraGGkDkrD6/q/IpP1KOLsJFQyBCRvkZQNvpPr8dQCw/cTMwwsSzIUKlRBlASgA+v4S6Lj8k1zHCUCIjQoH7GEBa7Dy/7HMvP3VsM8KllyRCYaoYQONrO79dITA/ig41wggTJkKighhAa4c5v7//Lz/BszbCQpgnQl5lGEC6Zze/fpsvP+NGOML6GSlCdkcYQB6aNb+hWS8/f9o5wiHBKkLhLRhAEhE0v0whLz/WcTvCvEEsQkJLGECULDO/llEuP1ABPcIKwS1Ck00YQLbfMb8lwy0/OJg+wr8/L0IyXhhA5X0xv6daLT/eL0DCFrkwQoaYGED4XTG//WgsPwzBQcLyLzJCt4EYQLOQMb+x1iw/A1JDwpKkM0I6lhhAC70xv9WXLD8r10TCfhI1QhOzGEDA1zG/PzEsP9xXRsJUfDZC/rUYQHoLMr9ROiw//9lHwuXkN0LK1xhAb1gyvyzUKz9kW0nCNk45Qrz4GECYAzO/cpYrP33GSsJDsjpC/gQZQNJ6M79KlSs/CV1Mwl4LPEIqOBlAZEU0v70bKz84v03Cvl49QhCDGUC7BDW//z8qP0oyT8L5vj5C0rAZQGGZNb/RxSk/cINQwgoTQEK2EBpA11Y2v/SVKD+m9lHCBl1BQglcGkBUaTe/8tYnP9ZjU8L5oUJCs7IaQARmOL8g4iY/JbBUwjfwQ0K6ChtATE85v87fJT8zD1bCIDZFQrJuG0DMqzq/q9gkPzRjV8LEdEZCcJ8bQEBQO7/LVSQ/FLZYwnixR0JA3RtA/H08v6vRIz8cAFrCdPVIQoIoHEDhaT2/aP8iPzBXW8JmKkpCq1QcQGk0Pr8PmyI/ZKNcwgZbS0KvXhxAxHE+v9eJIj+e6V3CWo1MQuJyHEAM7j6/PGciP/YyX8L7vU1CH5UcQL3TPr9t1iE/OF5gwrDwTkL/YRxA/mc+vzZ5Ij/mjGHCABtQQhI/HEBaKz6/D+0iP5jPYsJIRVFCLCwcQPS7Pb/6DiM/mvxjwjtrUkIr4xtApCU9vxj4Iz9gG2XCkpRTQrmkG0Dl0Dy/fs8kP8IvZsKFplRCw4YbQCgxPL+ZCiU/gmlnwujLVUI9RBtAB0U8v/IYJj+kl2jCld9WQlMJG0Duwzu/otEmP5aYacIG8ldCu+EaQLRKO7+hQCc/arxqwg8EWUKpqxpAuzU7v7wOKD80sWvCjgdaQm5/GkBm8Tq/+KMoP1TabMLlGltCE1QaQF8sO79hZik/OsxtwhopXEJFKRpAVaA6v5TaKT+e3W7C9zNdQiYcGkDhqzq/FhMqP2rZb8KjR15Cmg4aQCzeOr80XCo/btNwwqxSX0IR5RlA8iE6v5W4Kj/U6HHC8GJgQlXZGUDmJTq/sOgqPxDdcsKzamFCRLkZQHagOb9SNCs/huZzwiJ8YkIguRlAPCw5v9QHKz9gznTCaIBjQom+GUDRkTm/uxkrP+irdcLagGRC6LsZQOk+Ob8KBCs/bNB2wn+aZUI7sxlAx/U4vxIKKz/ouHfC26tmQtXEGUDy1zi/xbgqP4SkeMJMs2dC0NUZQABTOb8fpSo/aKZ5wtC1aEKQ6xlAGwg5vwoyKj8Sm3rCmMNpQsEdGkARtDm/hK0pP2Sfe8LeyWpCIG8aQCCMOr/6vSg/4Id8wjbJa0JIuhpAjzg7v+fVJz9UfX3C1s5sQvMGG0Bo3Du/PeQmPzZ+fsI74W1CcmAbQMnuPL+w6CU/9nd/wlXTbkIIwxtAYW4+v4PwJD+ZJIDCnMpvQssBHEAUYD+/sFAkP7zLgMK5x3BCp5ocQKgNQb/mjiI/tz6BwhTGcULf2RxA639Cv3EZIj/QqYHCTtdyQkBAHUAuSEO/SMogP3IbgsLMznNC3pkdQOSuRL9B5R8/HpiCwhvGdEJS6h1AqEBGv1AyHz9DFYPCRMt1QvEuHkAg60e/iLUePz+Og8KDu3ZC3KceQCF5Sb8rXB0/4giEwvqYd0JRIB9A/B9Kvwi0Gz94h4TCFnV4QlOLH0CoxEu/uJYaPwj/hMJgWHlCwtUfQGRaTL9Unxk//GuFwrY1ekJNHiBANhVNv3a7GD8643fCKaYev8o6sr5sN2k/FQqsvrQOdsJdIkq/jE+vvpS+aT8fWam+dDR0wlkOdr+11bC+T8NpPzrbqr48XnLC/KaQv7N5tL4O0Gk/+nauvuKOcMIqM6e/4MO6vg3QaT/+rbS+pLRuwjC/vr9c98C+RnppP2ixur4w2WzCEY/XvyJFx74HFWk/SsrAviIOa8IimPC/lQ7MvnHCaD8iasW+kDdpwm9hBcAugdG+PLNoP+vOyr6eY2fCOJwSwO2A1b4Do2g/HcTOvoKXZcIcICDA58LYvh+8aD8FD9K+ns1jwviVLcAK9Ny+1ihoPxX/1b6EBGLC5Tc7wN0x3r4ISmg/Y0vXvtRBYML2B0nAgXjhvhIoZz/dENq+0HxewsaeVsCk8eK+pmtmP+I0277Yt1zCWktlwKFg5b6jK2U/URHdvmL6WsLrynLAj2LmvsXnYz88ft2+GDdZwqxvgMDeFui+/gJjP9nG3r78f1fC9lGHwGev6r4I92E/d97gvn/OVcIuiY7AaqXsvtvSYD94R+K+YSBUwtVolcCQm+++1gxgP/3Y5L4HbVLC8K6cwJHO8r5xi18/lcbnvsG5UMKcqaPAoZP1vq5SXz8vauq+6xFPwviyqsDQXfi+gn9fP2xF7b6LaU3CHaOxwGak+r7riF8/cY3vvkrRS8K9hbjAJRD9vgH1Xz8CLfK+ZjZKwrtdv8ACVAC/HOBfP1G49b72okjCSxzGwCbQAb/wJ2A/ZNX4vrAkR8JWAc3AuB0Dv1hHYD/wgfu+cJlFwkGC08Da/gO/YHFgP6Fb/b7NF0TCaGzawGbyBL9eg2A/Rk7/vvWgQsKU2eDAVjMGvwhdYD+F3wC/qRhBwvRf58C+mAa/pIlgP4lRAb9VwT/C96/twIhhB78GH2A/2v4Bv4x2PsLCIfTAG+kIvxV8Xz+kXAO/3wI9wprk+sAefQm/6fpePynOA7/PuzvClocAwSoVCr91u14/rlUEv6xzOsKSugPBk7wKv61AXj8g3AS/TD45wiYLB8Fw0Qu//k1dP+SuBb8Y+TfCAjQKwa4JDL8xhlw/Y68Fv/3INsIjVA3BmZoMvyGAWz8Y9wW/yqo1wk2oEMHBUA6/DDhaPx1RB7/phTTCF8YTwS08D78wjFg/HsIHv+R3M8JS7BbBq5oQvxW8Vj/lmQi/RIgywuASGsFDIBK/DyFUP+BaCb95cDHC0E0dwU5YE7/Fa1E/v8MJv+9vMMJ8diDBNkAWv8cwTj+Rqgu/t44vwhR9I8Fw6hi/lDBKP2cQDb//ly7C1PYmweqJHL+ZEEY/xFEPv7WxLcIB/ynBpXggvwPuQD/gfxG/G9cswlA8LcHd9yO/kwE8PylDE7+PBCzC1L8wwU1bKb8s9TU//mMWv/wuK8LTRTTBRE4vv33mLz+E8Rm/UIUqwhCQN8GA3zW/llkpP0fJHb+8zinCv+U6wRmSPL8PQSI/sGghvx0lKcKFaz7BWTJEv6+SGz876iW/rIAowsQ5QsFUiky//9IUPzzqKr++zSfCEu1FweNXVb+e9A0/9yIwv3crJ8KwfknBgRRev/l7Bz+rUjW/Kp4mwi2rTcGejmi/DbYAP/zWO7/87CXCUHRRwYEOcb+1F/Y+Se5Av+Z+JcLkjlXBZxl7v0586T6KBUe/8xUlwr6pWcGwc4K/eKTdPjr7TL92niTCadddwd5jh79ys9I+DixTv14mJMLtcGLBQqeMv+VHxz7Iw1m/8skjwoivZsGnSZG/6A++PlLHX78adSPC6k5rwf18lr/X8bI+cTxmv/AMI8LCBHDBKyqbvzacqT5tOmy/c7wiwozFdMF4u5+/HsefPhHOcb80WiLCO6d5wb2so78cvZY+/2J2v20mIsIGxX7B1ASov28WjT6eg3u/XuEhwm7sgcGs86u/fBiEPrsDgL9qtyHCYF2Ewcdqr7+yZnY+DsuBvy5QIcK7A4fB/26yv/ROZz6MXYO/Jkkhwh6micESZbW/TuRUPo2ChL8TVSHCWCuMwdxCuL82s0Q+gMSFvzcFIcLgDI/B07G6vwYKNT7jmoa/ZBshwtDGkcGA27y/2nIkPr8Gh79+/CDCf3KUwZbTvr9QExQ+ET2Hv5oxIcImVZfBTRjBvy4tAT5vcYe/mkFawoCo+8FLrDfAVUlwv3ath76+2FvC1Cr8wUlNOMDkx3G/UBuDvv5lXcLyqfzB8nY4wDf0cr8MIIK+HPdewoAt/cGtcTjAvjh0v5iegr6cnWDCyrT9wbJSOMC+CXW/2suDvhYzYsLkT/7B2io4wKI3dr9kWYW+rtljwpDo/sFIGjjAe/x2vy4Shr56fmXCynP/wSdSOMBIfHe/BXWEvmIcZ8K6BADCt2U4wA2ad78/4IO+HsNowhFJAMLaozjAjct3v9L7gb64W2rCbaIAwjsQOcBrAXi/00x9vpD6a8Ji2gDC5pg5wJCmeL+aE3W+HJRtwhYdAcKdazrA/ht5vx4gaL40M2/CwlYBwnsCO8A18nm/CBNfvljNcMJwkQHCZP87wCHEer8TnU++4FhywhbBAcLj4TzA4KZ7v6fNQb5G5XPCXO4BwhbgPcAxtXy/6EkyvnRndcKZGgLCJss+wIh9fb8d1yO+QPR2wkE9AsIorz/A1C1+vybHFb5kbnjCaV4CwkuyQMDoBH+/FcgFvvD8ecJ4dwLChaVBwB74f78fi+29mGt7wqGLAsLHf0LAXk+Av7py0r3y+nzCtaECwt17Q8CDkIC/eAGzvUJsfsJosQLCpVhEwMbWgL/Kdpe9YOt/wnqvAsIyGUXAFwOBv8q/fr0XrYDCYK8Cwhz+RcAeJoG/EltFvdFxgcK1qgLCtc5GwMQqgb+o9RC91iqCwtGlAsJfk0fAMTyBvwArv7wR5ILC8IYCwjllSMADH4G/OmErvHiUg8J9YwLCGMlIQORAgb/nNY47wVCEwvBFAsKmA0hAtBiBv1KuhjwOCYXCWRYCwisRR0DnD4G/JzYAPeK5hcIH5AHCjh9GQPT+gL/a1jw9QHGGwkqxAcJiO0VAYO+AvxkWdj1lIYfCXHUBwmUhRECpg4C/6TOePSzQh8KkHAHC3B5DQEZfgL+Zib49aH6Iwp/aAMIrHEJALdR/v+aa3j0cJonC6IAAws3yQEDW/X6/3LsBPi/FicJtIADCpsg/QGqAfb8//BM+DGuKwnpq/8E2ej5AIyN8v4V1KD64EovC4Yn+wfcwPUBNy3m/Eis8Ph2li8Les/3BlMc7QOdOd7+mqVE+DTmMwojI/MHpaTpAUpd0v5cnZj6m2IzCirb7wRP3OEChynG/3rp7Po1mjcIGlPrBPlw3QP7Ybr96v4k+zfWNwrRi+cEHxjVAEM5rv1RalT58iI7CRi34wRUdNECsGmm/hImhPuAKj8IS9fbB0VAyQL51Zb92YK4+XaCPwkyc9cFCszBAJJFjv2xTuj7+IZDCw0H0wRffLkDroGC/PHrHPmyfkMI3xvLBTvssQPfYXb9CFtU+thaRwvA/8cFTEStAqCVbv5fZ4j5KlZHCrL3vwcgzKUB8lFi/RDzwPkoYksLbEe7BEmMnQB1DVr9FT/0+J3iSwuRe7MHsfyVAiJ5TvzJeBT+d8ZLCGrXqwY+CI0CPuVC/sWMMP9Bak8Km6ejBO5whQGnjTb8sDBM/1MKTwoQS58H/sx9A3J9KvyKSGT8DMJTCIi/lwezgHUBZYke/Yb0fPwiOlMIyQuPBa/cbQFASRL/ZMyY/KOuUwihA4cEaLxpA8JhAv7gNLD++S5XCiFjfwXRFGEC71Ty/rUcyP/SulcJVRd3BP7cWQAiQOb8rOjc/zgiWwg/72sEx/xRAGsA1v0uTPD/QV5bCYADZwRtvE0DkOzK/OGFBP0K4lsIRx9bB++8RQIUQLr+UlEU/8AOXwpqL1MF5bhBAb3Mqv5EHSj+0T5fCMFzSwRQ2D0Ay2Sa/DUhNP7Cbl8KTD9DB5uoNQJ9sI78d4lA/+wKYwkrRzcFGBQ1AYoEgv8AVUz/hRJjCNI3LwZ0HDECwWx2/EodVP02WmMLfGsnBohoLQMcaGr9Qnlc/uMeYwlz3xsHmeApApj0Xv2CvWD8JF5nCiKfEwWXZCUBs9xS/xwJaPwVimcK7TsLB6VgJQBrnEr8b8Vo/p5uZwtbcv8Hf8ghAHMkQvyFpWz+u85nCBaq9wbHeCECNSg+/i+laP3A7msLyX7vBA9EIQH6iDb9VOFo/qIKawmACucFC0AhAuCsMv7RtWT8ur5rCbqq2weDlCEAXlgq/ZzhYP2KAYsK+UP3BoqdEQKSpgL/aco09nkBkwiDg/MFvQ0RASmyBv4t6mj3g9GXCzXD8wd88RECoy4G/h4ebPRimZ8K0CfzBMDVEQA8egr+FsZw9TGhpwkmg+8GiQURAkBCCv0gXmz0UFWvCMEv7wVNiREBbHIK/FP6WPebLbMIm9PrBOH1EQHLSgb/zbpM91IBuwgCV+sGwV0RAvFOBv17flz2KLHDCQz76wZxXREACkYC/426XPcDdccKS4vnBa0dEQL3Jf7/9D5k95nxzwhCc+cFOGURAHXV+v3Nrnj0uJHXCGDD5wSfeQ0DYt32/E5ClPXTAdsKX2fjBLXFDQP60fL8Fy7I9+l94wr94+MEUSUNAJVx8v1Kotz0QAHrCvh74wZXWQkChWXy/KuXFPeyMe8KmuvfBp3tCQOiEfL9CR9E9UBd9woFY98FdGEJAvRN9vwXi3T0ql37C0/f2wdS+QUD4fX2/VTzpPYEOgMJdjPbB8HBBQLTofb/3J/M9g8iAwmct9sFFAUFAYIx+v1G3AD5RiIHC7Lz1wYagQECyJ3+/TO4GPlM4gsK6VvXBV1BAQCpsf7+gBww+ZPeCwhz09MFI6T9AgoB/v56EEj7UpoPCZpD0wbicP0BZjX+/u1YXPg1ehMI0FPTBzXo/QFtgf791ahk+XAaFwnei88EbKD9AuwF/v0h+Hj4cv4XCizbzwUvnPkAVpH6/2nEiPvdshsKZzfLB/LQ+QHxbfr9RgiU+0haHwlo38sGidj5APtJ9v9g9KT7AuYfC1KjxwdwiPkAS+H2//IouPvljiMJgKfHBFuo9QMx8fb9w7zE+jQ6JwkGR8MG4jT1APoB9v9+6Nz6KrYnCb/TvwcQ1PUBZc32/Ojo9PmhSisI8Xu/Bxuo8QFtdfb9I5kE+JvGKwtfB7sEHYTxAP8l8v8FQSj7yjovCBPPtwa3zO0C/pHy/Kh5RPkIpjMLCXO3B04Q7QEPze78FyFc+HsCMwmCV7MEy9DpApmV7vwucYD5gSo3CEcXrwRlfOkA+Mnq/mWhpPhPgjcKs8erB86U5QOkkeb/tf3Q+S3GOwuAX6sGy7jhAn053v8wKfz608I7CKEDpwWkKOEAPOXW/YBiGPsh0j8LKaOjBiUE3QNIGc7+ZvYs+lv2Pwmhr58FrVzZAHc9wv31dkj6GfJDCUWTmwb5KNUBgY26/CvKZPjH2kMKvXOXBmkk0QELna7/GFqE+g3aRwttP5MHxIDNACqFpvxB2qT7Y45HCV0LjwQbfMUCMt2a/5FGyPg5lksI0GuLBULwwQOwNZb/5nro+qNKSwh3x4MGMXi9AvcNivyhzxD5mQZPCgq3fwV30LUAUoWC/uqvOPkGnk8J/XN7B/XUsQMRpXr/nbdk+dRWUwqUO3cH7+ypAz1Jcv1IR5D5EgpTCMqPbwVqNKUB2glq/3XTuPo/XlMKMO9rBpwMoQBNyWL+bhvk+mkOVwh3J2MGEVyZAdwlWvzm4Aj++mJXCv0TXwf++JECbvlO//GIIPy3wlcI/udXBmyIjQDLqUL8N7w0/AFSWwvod1MEyiiFA8FROvw94Ez9SoZbCdHPSwQLXH0AMbEu/zUoZP5/ylsK1w9DBTTweQJicSL/jvR4/pUWXwjstz8GwiRxA6kdFv6RaJD/am5fCmWDNwaAQG0A7b0K/HTYpPwPrl8LwccvBIXgZQKgIP7/QUi4/GDCYwtS6ycGJABhAtOs7v4f/Mj/jhJjCmNfHwZGPFkAgGji/Bz83P67FmMIl8MXBDBoVQFvnNL9FyTs/CgiZwqAUxMHt5RNANr4xv4tHPz+URpnCTynCweenEkAmly6/DedCP36rmcJ2P8DBSbwRQN4GLL9tdkU/2uSZwhhAvsHOuhBAZ00pvxBGSD/5LprCpim8wQDSD0C4ZCa/KpVKPwhUmsLgU7rBXh8PQKbEI79KJUw/fp6awvBVuMEjdQ5A48khv6DfTT8K4JrCWkq2wYrsDUD9FCC/UjFPPw8Ym8ILQbTBPoQNQG94Hr8HClA/3GabwtpTssHaXA1A9yIdv+T9Tz9YpJvCXXOwwY1LDUBX1hu/OZxPP1jjm8KObK7B1kMNQAzjGr+iQE8/5AqcwuptrMEcSw1AL4cZv/1zTj/shmTCVg/9wcvRQECypoC/s2UEPl5NZsIecfzBmYVAQClpgb8hmQk+3AZowprU+8FHkUBAJr+Bvx8KCT6AvGnC6ED7wSiPQEADA4K/mVAJPqSCa8KlqfrBgZtAQCbigb8zdwg+zDJtwv4m+sHdukBAVtmBvwB3Bj7w627CdKL5wevUQEDoeoG/rKAEPjqjcMKKFvnBqa1AQGXngL+5zQY+HlFywnCS+MGtqkBAzhCAvw2NBj6aA3TCfgr4wTmaQEBtpX6/5DAHPhykdcLXlPfBB29AQEUufb/yfwk+uEx3wuj99sELO0BAQFJ8vy2CDD6I6XjCIHz2wULXP0CSK3u/GWYSPlKJesI88fXB57s/QK+4er9p+BM+ECt8wppt9cHuWz9Ao6V6v8HqGT5IuH3CruH0wQsTP0BgwXq/T3wePihDf8JXWPTB/MQ+QH9Ee7+DgiM+nmGAwlnR88EwgD5Ayqd7vzTtJz5RJIHCQEDzwdZHPkDGC3y/LJMrPmTegcJNvvLBB+09QC6sfL+BdzE+Z52Cwmgq8sGwoD1AwEB9v7ZzNj6qTIPCpqPxwRtjPUBKfn2/RWY6PkwLhMLMIPHBlRE9QAqNfb/wiD8+zrmEwiWe8MG71zxAH5B9v5YrQz5tcIXCpQTwwbDLPEBNW32/0thDPtwWhsLpdu/BUYw8QC/tfL/fp0c+gc6GwuLw7sFoXTxAlY18vx5zSj6Ue4fCPm7uwaM9PEDkPny/eVFMPq0jiMJ6v+3BjRE8QHK1e7/X3E4+fcWIwl4a7cFV0jtARuB7vzXkUj6ibYnCJIXswVCpO0AkZXu/okJVPtEWisIS2evB5F47QJJze78j8lk+KbSKwgko68EYGTtAc257v8JPXj4vV4vC/37qwfjfOkDlXnu/ut1hPk70i8Jd0enBp2U6QGnaer/uTGk+m5CMwr7y6MELCTpAc7x6v5sNbz4MKY3C4k3owUuqOUBTEnq/Ka10PjG+jcKrd+fBiSo5QE6Seb8scHw+PEaOwnqZ5sEqpThAFmd4v3QZgj6Y2o7C3rrlwTP8N0BLY3e/bx6HPo9pj8Kk1uTBZFU3QB2fdb/x2Is+HuePwj3z48GSfzZAVZdzvynvkT7kaZDCPxTjwUHINUCId3G/0ACXPozwkMLyDuLBIu80QC9Vb7/YE50+mG6RwpYB4cHi9DNAbPpsvz0PpD7q5ZHCZvbfwSsHM0B8j2q/EpSqPpRkksJY5d7BCfIxQA5ZaL/yVLI+AtCSwizV3cF3xTBAvotlv2GHuj5JT5PC8arcwYW2L0Dc6mO/6jbCPrK6k8I6gNvB9m4uQDK2Yb9DXss+eyiUwu072sE3HC1A5a5fv/rk1D7BjJTCa+nYwUSyK0BUjF2/NAvfPoz5lMKim9fB4U8qQDWNW7+E/Og+RGSVwi0y1sHs9yhA19RZvye38j5EuZXCUM3UwWuDJ0Dn3Fe/QCz9PjsklsJOW9PBAO4lQOKLVb+rNgQ/k3eWwiXb0cGzaiRAmV5TvzSVCT+NzZbCIlTQwf/iIkCTnlC/g9QOP+Ewl8Lou87BXF0hQNEmTr/KGxQ/g3yXwkEVzcH9vB9A0k9LvyqpGT/8zJfCzWvLwV0zHkDvnUi/FuIeP7wemMKi28nB4ZIcQElYRb+/OyQ/DnSYwogUyMHgJhtAqZVCvzPrKD+zwpjCCy/GwXWdGUD0Qj+/V9MtPy4HmcJGfcTBrjIYQO40PL/bUjI/AluZwpiiwsGSyxZAsG84v5BwNj8Nm5nCzMLAwSRhFUDWTjW/QtY6P9/cmcIQ8L7BMDUUQIg7Mr/3PT4/0BmawpkOvcEV/xJAjBkvv7fAQT/Gf5rC/i27wRAZEkD0mSy//kFEP524msL6NLnB+hsRQNLxKb/7CEc/YAKbwqsot8HXNxBAlBQnvxNMST9OJpvCNVu1wTaGD0AMfSS/2t1KP7Rwm8KHZrPB2d0OQKaPIr8zmEw/TbGbwi5iscGJVg5AyOogv57tTT8Y6pvCcWWvwV3vDUB/YR+/gsxOP104nMI3gK3BjsYNQEUVHr/zy04/UHWcwuKrq8GEtA1AH9ccv8Z1Tj+Gs5zCXK+pwa2rDUC++Ru/oSpOP0vbnMJruafBMK8NQF2nGr81c00/oFKCwhY9d77GIRlAWW09v/oMLz+MZ4LCsuROvqZBGUBuYj+/VFEvP4J8gsK0sCm+sakZQLigQL/sKS4/N5GCwgpgBb4h+xlA6aJBv0tFLT/BoYLC1GXAvf1DGkDaVEK/t2MsP76xgsKru4K9xp4aQCYnQ7+7RSs/97+Cwl3WDb26BxtAdQlEvxj0KT8BzoLCtCIDuQBNG0BeckS/RwQpP6DcgsLGDfI8DqobQMeyRL+zpSc/UuuCwkKCZz2AHhxASiNFvzf7JT+/+ILCQrWrPW+AHEA7PUW/2HskP1gCg8Lm4eQ9cPkcQH/oRb8K1SI/YQyDwmWHCj6mYR1AHLtFv0kkIT8OF4PCWfUfPrDmHUBGYEa/4EsfP6Yng8JgqDY+glQeQPr1Rr+Lyh0/0y+DwmJcRT7iwR5ArnJHv09CHD8TNoPCwD5bPoAYH0Boj0e/7PMaP7A+g8IWK3A+VWwfQJalR7/Nrhk/D0qDwifgfT6DsB9AjJpHv/ycGD+EUYPC+N6FPmThH0D44Ue/Y/MXP5NYg8Lcyow+UQogQJ0QSL/6YBc/xl6DwhZSkT6DNSBAzlxIv13PFj8UZYPCy9mUPmdSIEAYpki/TnUWP+Brg8JYPZg+3WogQE0vSb+rQRY/5G2DwtFFnD65jiBAxNJJv1jpFT8ZboPCryKgPiWDIEAI9Em/GiIWP7t1g8Lxl6I+bn8gQMxBSr9SShY/dHmDwuBdoj5TcCBAEF1KvzCPFj/LfIPCQ/CkPntZIEBvSUq/YOMWPxR+g8IAtac+Iz8gQOxRSr++Thc/b36DwkiwqT42DSBAUuBJv2bvFz8qgIPCTKWsPrLuH0DIw0m/Hl8YP8OFg8JUgK0+TLMfQMReSb9lKRk/D4ODwkjqrj55eR9AWVtJvywOGj+Oh4PCDO+wPpQ1H0At8Ei/H/gaPw6Ig8Kc1rQ+i/seQE3+SL8X5Bs/04mDwlbztT77mR5Ax6RIv4BKHT90kYPCyZi7PuBaHkC+iUi/Iz0ePx2Sg8Kj3r4+a+0dQEiISL9z8h8/bJODwsHHwT5upx1AZz5Iv9rwID/nlIPCWUjHPtdAHUA8C0i/QnoiPzGZg8Ic3cs+btMcQBa1R78NEyQ/3pyDws5gyj5xehxAIsJGvxMhJT/InoPC4VvQPtIfHEDIZ0a/wWwmPwqig8Lq2dQ+DKgbQJSpRb8xCSg/NqWDwhxk1j4sVhtA6N1Ev3kHKT9uqoPC0KPZPm8BG0DwUUS/uSgqP2Gqg8JAa9s+BaIaQODDQ78fdCs/Gq6DwoaR3z4IWxpA7y9Dv3haLD+Es4PCKHLjPtX/GUCNzkK/vaUtP1C2g8Kg3uY+xbYZQMApQr+WjS4/kLODwnV47z77bBlAAqdBv1WFLz+GtoPC52LwPrIvGUBlLkG/RE4wP2q9g8J4HfY+y/cYQAr6QL8qHDE/vr6Dwvl49j4Z3RhAHOxAv+aCMT9lxoPCNbD7PnOjGECXxkC/9V0yP0DIg8K2XAA/ookYQA1zQL+GpTI/MsyDwvovBT/2VxhAbClAv9RRMz/I1YPCEJgHP5BWGEAnYkC/NW4zP5ffg8ITMws/RDYYQHh/QL9H/TM/yeCDwuHECT9kIhhAFG9Av5BHND9m64PCmEYNP6sGGECaYkC/ZLM0P4Lzg8LOwA8/YwQYQBp8QL/1xjQ/mACEwmzgEj9U+hdAzr5Av9cKNT/kC4TChhYXP4n5F0BK4kC/ZBw1P5AZhMKk7hs/9/oXQFMIQb/nJTU/FSSEwgFTHj9yCRhAUQJBv3DoND+IL4TCVh0hP7MSGEAVN0G/8dc0P+syhMJ25iM/wR4YQJTdQL/OgjQ//0WEwk0JKD8CIxhAar5Av/JkND+JTYTCFAkvP3MtGEBMLkC/jAA0P+5YhMLG3S8/JkgYQEXnP7+NdzM/lGOEwtZbND8rURhAa0o/vyMUMz9TcoTCZ0M5P2JXGEABYz6/VZ4yP7p/hMKwrz0/+WYYQP2dPb+OEDI/IomEwtArQD/JhRhA5ic9vyZlMT91nYTC6PFCP/GcGEBpMTy/1KUwP8OshMJpzEM/qbsYQJB0O7+E3y8/db2Ewi4bSD9syxhAHMw6v55dLz8rzYTCZapLP8DlGEAe/Tm/f6IuPw6/g8JnFFS/L47uPsj/WT8K7+A+iuGCwijOJr92V+8+HtRdP7WG4z76BYLCohnzvhQN7j767l8/JkDjPn4xgcKf1Je+zi3sPgEPYT8v7eE+El6Awi7P8r2Jheg+pRFiP0nF3j4EE3/CduJwPaRI4z5rtmI/Tt/ZPgRufcL4aWk+atLdPjkQYz8an9Q+CMd7wmHpzj7mqtk+7rFjP9TJ0D5SNHrCPiQSPwES1T4uS2Q/CYDMPjSYeMIQjj0/WZnRPnt1ZD8wJMk+Ugl3wo9WbD89CdI+83ZkPz6TyT7AbnXCRFGMP2Cy0T6kMmQ/ICHJPorsc8K5w6E/yh7QPtayZD/1x8c+BmpywuCYtj9iM9A+aBZlP1YFyD5+6HDCOSLLP9Em0T5DRmU/pgnJPiRyb8KXMt8/uPnQPvTZZT81Gsk+fN5twqZc9D+Fl9E+k/ZlP0PCyT5+aWzC4nsEQHy70j7y42Y/fkbLPvr9asKu/g1ARRLUPqZWZz/3ysw+ioZpwqX3F0CTH9U+OzhnP9PJzT48ImjC7IohQNTV1D5kN2c/K4DNPrKzZsLuVytAiXbVPi90Zz+jOc4+ilFlwhK2NEAUldU+2pNnP2xlzj628GPCCMQ9QOkH1j5Ja2c/esbOPuiSYsJ+B0dAPpvWPnDUZj/lGM8+eDRhwvZBUEAvLdc+gL9mPwahzz4c8F/CbexYQNyP1j4xcGY/8uLOPvabXsJKdmFAtH7XPgpqZj+Lzc8+lmBdwmYuakBIw9c+MKFmP0Mp0D6WKFzCTftyQD2+1z62cWY/8Q/QPgLsWsK/lXtAsp/XPuSsZj/lCtA+YLFZwlCigUCyD9c+NHRmP51jzz4YeVjCKrCFQInp1j6lFmc/64LPPg5kV8KTm4lAgobWPt9zZz8VSM8+pEZWwnSqjUCiatU+IHlnP+Mvzj5lKFXCNqGRQFt21T5tjmc/j0TOPu4oVMLmpZVAgTjVPjuEZz+5As4+Ch5Tws+GmUBsotM+O6tnP0h/zD6aJFLCG1qdQNdt1D5+EGc/cgjNPtAsUcJcW6FArnLUPrO6Zj8j6cw+6D1Qwv99pUC+nNU+5rNmPzwOzj6qaE/CIwapQPg+1j5opGY/zKjOPtiNTsKa76xADQPYPpwVZj/fLNA+VM1NwpYOsUCLFto+b7RlP8wS0j5KBU3CwuC0QI/v3T4SNWU/o63VPopLTMJXk7lAE2fiPvzmZD/a+9k+hJ9LwsiuvUD+vec+rtVjP5TP3j7e6UrCTjXCQFmP7j4hh2I/S/zkPoRaSsJ2pcZAoVX1PsxdYT8+LOs+qsZJwrtuy0CxLf8+yQBfP/XM8z6GOEnCmv3PQMqMBD8vSVw/80T8Pgu8SMJ30dRA3xsKP4iVWD9EqQI/MUJIwtjR2UCQURA/jBpVPx/YBz+G20fCmEvfQIObFz9Zu08/XHkNPyWAR8JPAuRAcB8fP1/EST9XCBM/TEpHwliz6UCamSc/MHpCP8T2GD+AFEfCFSzvQHSPMD82uDo/BQYfP1XPRsIww/RA9TY6P3fgMT/pJSU/tLNGwsqC+kDEYEQ/blsoP0k+Kz+5nEbCjA8AQeczTz/Q0h0/aEExP+pqRsLlEANBUB5bP+L1ET/VWzc/O4FGwtMEBkEnGGc/0NQFP2L2PD9UfUbCyA0JQQF0dD88hvE+AO1CP7+bRsIlNAxBY1mBP0AV1T6gi0g/l5ZGwv2WD0GuyIg/fae4PhUdTj9Y6kbC8rwSQTptkD9ASJo+K81SP3o3R8K8HxZB4CuYP0tMeD60Alc/xWRHwsBdGUHvJKA/xtw8PnkGWz9ioEfCtGkcQeCyqD8z//09i9peP1cSSMIO5x9BZomxP7UvhD1miGI/LmBIwl9+I0FWULo/sOkAPDngZT9PtkjCI+QmQXLhwj8Y/j29spxoP1ohScJIyypBVn/LP6tnyL3UYms/E2hJwtopLkHXfdM/PjsUvt+MbT/9CUrCPe0xQXlZ2z/EkkG+uZtvP1VrSsLYZTVBUr/iP8S5bL764XA/KOpKwiIbOUGoAOk/dYiHvow4cj9Nd0vCxaA8QQPs7j/C2pe+gB5zPwIWTMLkVUBBJjr0P5REpr5Zy3M/sLJMwujkQ0FfBfk//huzviRPdD9D5oTCyf6IQPggRr/ujzE/LK8wv88ZhMK1M31AzvxEvwKpMj/CADC/REqDwkpRaEAgE0a/xZEzP9dyMb9+fILCHZdTQPDIRr/QLDQ/s2Yyv9a1gcI0Vz5AN0FIv1DGND+eHDS/pOaAwuS5KEBi3Ei/Cdw0P4/ANL9OGYDCFKoSQJ4KSr9XBDU/qf81v4KkfsIrQ/k/7NBJv90bNT+rzzW/XAt9wlNmzD8pTkq/GS81PyhVNr8mcXvCmu2fP73aSb+YOjU/SOY1v1bXecJkIGc/jj1Jv6IANT/eMDW/CkZ4wiJDDz/mx0i/iGc0P+h7NL82rXbCLqNYPs3CRr/AXjQ/BXUyv4oedcJsmAa+PO1FvztJMz+7LzG/koVzwjLS7b7OaES/FuMyPwWGL79a8XHC+slPv9DLQr/xAzI/W5Utv8hjcMKq+JG/Z0BBvzsQMT9osCu/qsduws0lvL8lOT+/ZNgwPxWfKb8kMW3CkPzlv1+ePb8qmjA/0fYnv1yma8Jl4QfArN87vzijMD+LSCa/whhqwk/jG8BLSjq/mTcxP174JL+AiWjCXEkwwJSSOL/kZjI/PMEjvxj4ZsJSZETAKcU2vw4eND9+piK/CmJlwlMPWMDT+TS/x3w2P0HJIb/w2WPCPz9rwBvZMr/stjg/uIcgv7hNYsJAP37AbSIxv19/Oz9Z3R+/JMZgwtY0iMDs6S+/x6s9P9ByH7/yP1/C6RKRwOHULr9ZH0A/B0IfvxLGXcIyIZrAIqgtv1M+Qj9e2R6/kElcwofHosC1eiy/7A5EP0JSHr8U0FrCbkqrwEL2Kr+WB0Y/r4Adv5JXWcJ9eLPAVrEpv8urRz/Czhy//dlXwuqBu8DuJSi/GUJJPwnQG7/McFbC4jrDwEa8Jr+zl0o/PNsav2IaVcKC8srANdQlvzGUSz+TSBq/PqlTwirR0sDNDiS/1tpMP6bwGL9kTlLCnwzawDUnIr8jX04/SYkXv3r3UMK1ReHABNofv3S3Tz8MrBW/aqpPwjCq6MACrh2/Dg1RP27tE79gSU7CPHTvwJzuGr8dAVI/RXsRv8T8TMLcFPbATqIXv1qQUz9sqw6/MbtLwu7r/MBWkRW/v8ZUP2D5DL+LfErCFJ0BwYOiEr9xy1U/0lkKv3Q8ScLmrwTBAdUPv4JvVj/Zvge/gBJIwnG5B8HbwQy/hKxWP+/BBL/+y0bCe7wKwXJ+Cb83Hlc/n6UBvwOfRcL4fw3BRHoHvy68Vj9bGf++D4FEwnwyEMGC+wS/oEdWP0bw+b4aVUPCYEETwQKXA7+9SlU/s672vjU2QsIH2hXBmukBv6ftUz/RsPK+qiRBwsBXGMFxIgC/NqxSP16V7r5nCUDCxi4bwbzX/r5CxVA/azvsvuzwPsJ38x3BYjz/vlysTj9ci+u+4Pg9wnBqIMHjLwC//qBMPyOY67566DzCD6wiwS66AL+I30k/6TbrvqzgO8LbTCXBqQYCvxa9Rz+Nley+POk6whPzJ8EK1wO/dZ9FP37r7r7E3jnC1XoqweACBr8SlkM/kPLxviTeOMI65CzBp5AIv+YiQj9r/fW+APc3wpKuL8FvQQy/CldAP9D3+76B4TbCtxcywd4DD7/hxz8/Tm8Av38FNsKMyTTBUdISv9qQPj/qrwO/ICo1wllSN8HiSRa/cBw+P+DXBr/3OjTC+w06wRfgGr80mD0/Pg4LvwBWM8JL9zzBulsfvztdPT9cRg+/sHUywtqzP8FTsSO/q789P/aSE7+kpDHCbnpCwfeyKL+Bbj0/XlEYv0jMMMJkdkXBhHEtvxe1PT84Ch2/jOkvwg0+SMFSPzG/Ygw+Py7lIL+1CS/C9XVLwaS3Nb+i9T0/+EYlv9lPLsIpqU7B4kY6v3e8PT9kuSm/K4AtwpgJUsGicj6/S1A9P9y6Lb/ktyzCPi5VwXNXQr9Sdz0/XLQxv/DGK8Lch1jBkHhFvyaPPT/z5zS/Oh4rwkLhW8FAzEi/EqE8Px/nN79WiSrC1h1fwVJaTL+m/js/0kI7v8+yKcJqE2PBUcNPv7+qOj8FLz6/4Akpwt2UZsGvaVK/QEM5P65KQL+kTijCwuVpwSwfVb84fzc/rEtCv+zRJ8KZyW3B8+FXv3g5NT+7HUS/9HGJwhVq0b/DiBhAmUs8v2EBMT/miYnCntLLv6+mGECxRz6/+FIxP7yhicJopMa/rg8ZQGh3P78wIzA/MrmJwh2Zwb+6YBlAC2dAv+k5Lz96zInCc3C8v7OoGUAIBkG/Q1UuP/DeicITDbi/tgAaQLDCQb8EOy0/Z++Jwg7Ss7++ZhpATY9Cv+HtKz/W/4nCzeauv1ClGkC14UK/AxErP8EQisLNl6q/wP4aQJENQ79/uSk/piGKwkqdpr/cbRtAIm5Dvw4fKD8rMYrC8Imivw3IG0AKg0O/8rwmP7U8isKOZp6/IjwcQHojRL/UJiU/pEiKwoTkmr+RnhxAfO9Dv12KIz8rVYrCgLKXvxUgHUCOh0S/SLwhPxdoisLEXZS/AYsdQM8VRb/MRCA/6HGKwugEkr+p+B1A5o5Fv/e6Hj/jeYrCztiOv1pPHkDBqUW/zGsdPyeEisKDz4u/H6QeQMKzRb/sHhw/KJGKwhmiib+V6h5ATqNFvy8CGz8zmorCN4GHvz4cH0AY4UW/v1IaP/SiisLvZIW/gkgfQFUFRr/7rxk/XqqKwsrfg791dh9AB0VGv+EPGT9isorCr6aCvwyWH0B2gEa/CacYP126isKzdIG/HrIfQOb5Rr/UYBg/+r2KwvsNgL8q3B9ATJNHv9XtFz9Iv4rCcpB9vzPVH0CPo0e/zw4YP3nIisJOunu/a9UfQBTfR7/PIRg/vs2Kwto4e7/txx9AdedHvwJaGD810orC8jJ5v7myH0B5vUe/7p8YPxTVisLCEXe/ApcfQMi3R7/LCxk/gtaKwoCIdb++Yh9A1ChHv7eqGT9Z2YrCT1dzvzxBH0Da+Ua/kR8aPz3gisISFnK/aAEfQG2ARr9D8xo/xt6KwiPDcL+YwB5Am2tGvyrtGz915IrCXBdvv/tzHkB060W/H/EcP1PmisJ0VWy/Ni8eQGjuRb9gAx4/XumKwrgea789wx1AmoFFv+aKHz9e8orCiIhnv894HUAdW0W/1qUgP9zzisLdEGW/Xv0cQCxURb8ekCI/S/aKwnzaYr/zqBxAbwBFv5PDIz/U+IrCqDpfv28zHEC/xES/noQlP+H9isKuWVy/5bUbQHpdRL+IVic/0AKLwiVtXL86UBtAHGhDv9mTKD/ABYvCqJ1Yv77kGkADA0O/Kx4qPyAKi8J8ilW/nF4aQKRBQr8F8Ss/BQ6LwmrTU79V/RlAAmhBv7MlLT+QFIvC0KJRv5eaGUDq0kC/W3ouP/gUi8Jr3E+/+SwZQFc/QL9x+y8/sBmLwiUmTb/o2xhAEqY/v88GMT8uIIvC9jpKv3d1GEDWOz+/NnsyP0Uki8Le2Ue/oyQYQICXPr/egDM/lyGLwuijQr/2zhdAxA4+vxOlND/6JYvCxV5Bv6iKF0Dgjz2/r4Y1P7Eti8Jw0T2/5E0XQHdVPb+WZTY/2y+LwmzgPL/jLRdAjUE9v3zfNj+8OIvCmF85v33xFkCUFj2/eMM3P5I7i8L9Dja/utMWQHO2PL/YFDg/UkCLwuyRML/8nhZASGA8v8zHOD+aS4vCoFgtv9KdFkA5jjy/pd84P3ZWi8Km4ii/6nsWQMKaPL8Rbzk/hViLwiqWKb8KZxZAdXs8vx+3OT9GZIvCpC4lv6lLFkB0Xzy/Gxs6PwJui8JYACK/6UkWQCFmPL8NJTo/VnyLwu4aHr+HQBZAUpc8v/hfOj/SiYvCDwEZv7lBFkBpoDy/5l46P62Yi8L3gxO/XUQWQG6yPL+rWzo/1qSLwkgXEL+cVBZA0pU8v1sNOj8HsovC3pQMv9hgFkCktjy/Iuk5P1i2i8LqxQi/JW0WQP1BPL87hjk/KMyLwvnqA7+9dRZAmBA8v5hOOT9k1YvC4vj3vvaBFkBzcDu/Dto4P8Lii8KMJPS+154WQOgQO7/xPDg/Xe6Lwp2M6b7MqBZAR1o6v6LINz92/4vCpdHdvgeyFkCxajm/yD83P/sNjMJHf9O+98IWQI6dOL9JpjY/LBmMwmJFzL655RZAuhw4v+TkNT9qL4zCsE3FvlIAF0AJHze/QBE1P6FBjMKxNsK+VyIXQLhgNr+DOjQ/l1OMwmq3t75WNBdA9rY1v8isMz9AZIzCAHquvhVSF0CE4DS/K94yPyewTsKBbTVCcg7iv0JpOb68LH+/1/lOwh40NELppOG/hTYvvkG3gL/BRU/CcgEzQi6a4b/7Ay2+iwGBvxiTT8Lk1jFCvMfhv/XrML6LmIC/KuBPwo6xMELV3eG/IcQ1vozlf7/EKlDCTo0vQtM24r/V0j6+tNh9vwVqUMLkay5CwbDiv+zZRr5OW3y/rrVQwppQLUItz+K/u3lOvg5Der9UAFHCVDcsQm6p478c51a+aV95v8hLUcJCHitCRDvkvxoxXr7WQHi/vJBRwsUXKkI2teS/ATFlvmsGd7/t11HC2v4oQuPm5b/tWG2+vdZ2v/kbUsIn5CdCOGDmv0jDc74KwnW/3GpSwvLRJkKIjOe/BSZ9vo8Zdb+XslLCeMIlQs5p6L+/sIG+n9B0v+v6UsK4qSRCkdHov1fnhL4fjXO/BjxTwgadI0Lpk+i/ur+Gvlnlcb+ahVPCn5UiQjl36L9U+Ye+lORwv8LNU8KbfSFCH1Pov0kcib5n5G+/iBxUws55IEIXcui/kp+KvjAob7/kZ1TCamgfQvdq6L/ivou+emJuv/KoVMJpZh5C9Kzovygljb5p+22/F+xUwjRfHUIZP+m/hUGPvhS4bb8IMlXCgVscQmH76b9Ut5G+AoptvxqCVcIOXhtCJEHqvzTvk776nmy/KchVwg1gGkL+0+q/tmWWvvQcbL8qFlbC62gZQmqF67+6wpi+5ONrv+BNVsKJaxhCWkjsvya3mr5kD2y/YJ1Wwrh5F0JZf+y/ZqCbvoXea7+A3lbCMYAWQkni7L/N9Zy+w7lrv/0gV8KKkhVCpPjsv68tnb58v2u/IVxXwiOhFEITwey/+VKcvinma784llfCgK4TQkOy7L/W9pu+2wZsv+buV8JCwBJC6dbsv0Jzm76EpGy/1DhYwurYEUK2uOy/iHOavgkUbb8Wd1jC0uIQQmlO7L+e05i+iVptv0K/WMIs9Q9Cnrjrv/qPlr4BuG2/HhJZwkoOD0KWL+u/d/uUvoK5bb8WU1nC7CAOQtGV6r+QwJK+agZuv6iSWcI5Og1CaXXqv8BWkb6ds26/cMxZwi1NDEKDhum/Yi+Pvu1Mbr+0EVrCMGcLQsWJ6b9T9I2+xx5vv/Q2WsJRfQpCsLTovz7Ai75W726/Lo5awhaeCUJFdui/7uaKvoYCb7/I3VrCZ7QIQn5H6L//pYq+GtJuvzYFW8InzAdCtoHnv+aHib4rDW6/BERbwuHmBkJLjee/ECiJvkhgbr8yjlvCB/sFQhI8579+GIm+vs5tv0S/W8IfEgVCN+/mvyj2iL7YUW2/sOhbwvg2BELPvua/3WuIvg9Nbb8UJFzC5E4DQukz5r/w6Ye+6ZZsv2BaXMIZYQJCB+/lv3FNh76gdmy/HHZcwil2AULds+W/oNyGvkxNbL+6tlzCm5cAQnax5b8okoa+aHdsv9LrXMKxe/9B+Fflv/EGhr4fJWy/dhhdwseo/UH9BeW/J2SFvs3va7/ON13Cscb7QRy75L+4OIS+yBxsv4huXcI0IvpBIVzkv4Dvgr7LNWy/qIldwm9a+EF3NuS/OrGBvsmzbL8+qV3CPmn2QSgc5L/zF4C+PH9tvxC6XcK4rvRB67rjv8pKfL6J+m2/rPBdwpLI8kGc7+O/r4J5vro6b79iIl7Cg/3wQd4H5L9dAXe+Wi9wv05IXsJ8Te9B7Bbkv/oFdL5VOHG/vFZewvRn7UG/NuS/hI1xvn85cr+4gV7CbqXrQT6Q5L/LoW2+oh50v5iGXsJ48OlBI8rkv1n9ab7EsXW/UJpewuUc6EH1DOW/VIRmvlNKd79OuV7CMizmQUYr5b+QpmK+HLt4v5i8XsLmcuRBfWjlvxTCX76SHHq/YNFewjGm4kHB0uW/R7BcvkHoe788717CtsjgQfT55b+Yg1m+nTZ9v9wEX8KOAN9B3xjmv9hPVr4Jd36/Yvhewjk/3UHYOea/SHhSvsfvf7/iI1/CP5rbQTB+5r/ufFC+GY6Av54eX8I8wNlBmJ3mvwgzTb5cM4G/YjZfwsTh10GAKee/8vpLvkf2gb+IRF/C1SvWQdJR578OPUu+iz6Cv6pVX8LxZ9RBzonnv57ESL693YK/dlVfwgas0kGM0ee/Mq1Ivg0tg7+VelfC/Bs0Quqy+D4c2lg/7U7qPhZtV8IWDzRCl3z4PiMnWj+Ovuo+wG9Xwo/+M0KEgvc+pAZbPwg36j4ceVfCaPAzQjrA9j77w1s/6dTpPkx8V8Ib5TNC5kT3PjZ4XD9usOo+2ntXwjnaM0JE5/Y+rAxdPwid6j56elfCu9AzQuxR+D64fV0/ajvsPo55V8J/xzNCM135PqopXj9vme0+R4JXwvu/M0JOQ/o+46VeP5u77j58hFfC/7UzQlTZ+j7oEV8/nIbvPsmCV8LnwzNCHi/8Pq40Xz//6/A+UIVXwlS+M0IcTP0+UlVfPycY8j6PjlfCbLQzQqD+/T7AeF8/6dvyPmeUV8LrqzNCRgj/PldaXz891fM+g59XwhqnM0J5EAA/5UFfP5jg9D7FqFfC1pozQmhmAD85EV8/EXP1Pu2sV8KxljNCPJ0AP5jhXj/3x/U+7K5XwteSM0KelQA/BrFeP9uf9T6ds1fC7ogzQoDjAD+mq14/Vjj2Pje6V8LshDNCYcsAPxVoXj+H5fU+T7lXwm+BM0LmzwA/IQNeP6G69T4wvFfCpX0zQh2eAD/dil0/0Bn1Pqe2V8J3djNCs0IAP04oXT+mMfQ+GbVXwthuM0LD0P8+gaRcPw878z4us1fCYGozQvnM/j7KLFw/sfzxPvutV8IgYzNCr3n9PtWYWz8ZYvA+V6hXwoliM0LS1vw+PgVbP8d27z70n1fCKFozQuVO+z6rYFo/w6HtPsKhV8KyVTNCdnv5PuYjWj+At+s+ZJhXwglTM0Ibg/g+TElZPx5X6j5gklfC1U8zQqse9z4mNlk/IfDoPtSQV8JVSzNCQQH2PoV+WD93fuc+JIlXwuJKM0JwRvQ+kXFYPyzH5T5vhVfCVEYzQsoZ8z5/8lc/72PkPut+V8JcPTNCZ4rxPif4Vz8Z4eI+X3hXwr49M0JqD/E+1ZJXP5Y44j5ccVfCHjgzQpDi7z46b1c/ogLhPsp0V8K8NzNCGUnvPjV5Vz8NcuA+VWlXwg43M0LRWu8+/GBXP8J34D5qaVfCPTIzQrd87j4AV1c//prfPsZjV8JMMzNCjJ/uPjOKVz8r1d8+lVpXwkgwM0Ltcu4+DatXP1e53z4aVVfCMykzQix/7j5VAFg/t+3fPodQV8JbJzNCxZzuPl8WWD8AFeA+Ak9Xwk8kM0KZte8+j3ZYP3dU4T5PUFfCOxwzQhBF8D5/dlg/XuDhPopGV8KKGDNCkwXxPqUjWT8/7+I+FD5XwgYbM0JqdPI+tzpZP/5g5D72MlfCIRIzQnev8z6Yw1k/o9flPj4pV8LhETNCphP2PskqWj9ZYug+HC1Xwo4KM0Ludfc+CKNaP4v56T6oG1fClQMzQsUa+T6WUls/re7rPmENV8IaATNCWEz7Pl0VXD//ee4+TAtXwq4CM0LG2f0+A21cP3ks8T5CDFfC7vUyQmqm/z7uelw/3/vyPpAOV8KC8DJCjCEBPwDjXD+vyPU+0O5WwhH0MkK3iAI/+y9dP9K6+D5k6FbCX/EyQsvjAz+GX10/QYf7PlDmVsLZ5jJClD4FP1SPXT+9VP4+rOxWwhbuMkLBCQc/wkhdP2fiAD/c5VbCMuwyQn7ECD8Y71w/FYUCPyrbVsKG4jJC1P0JPw1lXD/lmAM/PtlWwireMkLGuws/TxBcPzFABT8221bCPuEyQuV3DT9+/Vo//a8GP4XWVsJ/4TJCexQPP3I+Wj8TFwg/+dlWwlbaMkJ8nBA/XARZP0NFCT/Nz1bCw9QyQocJEj/pL1g/6nQKP8/XVsIj0DJC74cTP7xxVj/+bws/+MxWwi/JMkIY2BQ/XzVVP5xhDD/y1lbCQsUyQp0MFj9mjVM/7xUNP6rcVsLcwTJCGDYXP0XwUT/YwA0/IuFWwkrAMkIZHhg/toNQPwY4Dj9j7VbCIscyQsFFGT8FzU4/NtYOP5z3VsLVvTJCwMsZP0hUTT835g4/HQZXwjvAMkJiUho/VttLPwX2Dj/nE1fCIsIyQm34Gj923Ek/GfoOP4AqV8JptjJCUkMbP7d8SD+P1Q4/skBXwgaxMkJtaxs/5+dGPwR+Dj8sVFfCv6wyQlddGz+4iUU/oAIOP9JaV8LNsTJCiX0bP6VURD8WwQ0/N4yMwoLUvUHOBRPAFG4rv4L/P78e14zCZV+8QWcyEsCrDim/yjxCv9AhjcJ7/LpBAc8RwAowKL/PZUO/Zm2NwgysuUFcvxHAWT8ov7mrQ79jtY3CrGe4QekGEsCcRyi/NZBCv8z7jcIrI7dBqIwSwDUtKb/M4EC/9DqOwjDktUGLABPACvYpv6RrP79Mf47CGrG0QVyUE8BSziq/nH09v5O/jsL6f7NBq/0TwNsYLL9Vazy/KgOPwqVOskGEXxTAcRstvxdWO7+gPI/CYUaxQSi8FMAUBy6/3Eo6v/t6j8LXGLBBkOwUwC71Lr+T8Dm/KbSPwkborkEsJhXAhhYvv/4YOb/m8o/C67+tQfJBFcAwvi+/EfI4v40ukMJmnqxBTC4VwJgmL796/zi/6GiQwk1nq0ECKRXAkLwuvxHnOL8kn5DCllOqQbw2FcCXui2/wkE4v2vXkMJjOKlBTw8VwNRlLL/wSzi/2A+Rwh0CqEHy5xTASFArv1lwOL80SJHC+ummQYu2FMAqYSq/fsw4vxKAkcJHvaVBVoQUwJ6IKb8GNTm/RLORwhqjpEFVQRTAjyApv64QOr+y45HChoCjQaUHFMC6Cym//Os6v74YksJJW6JBKdATwBg4Kb+g2zu/FU+SwjxLoUHMvhPAgFwpv7AwPL+mfpLC0DWgQQ2KE8ColSm/1Bs9vzm2ksJOIZ9Bf18TwGb0Kb+a7z2/Mt2Swh4NnkGfFxPAbPQpv6MOP793E5PCv/acQRPMEsA3sym/gB9Av9JAk8Ka5ptB644SwOyTKb9VBkG/03GTws3VmkG9LxLAZN0ov/UwQr9wmJPC1MyZQXrKEcAAyye/bUlDv5zFk8LRtZhBJ3URwJf7Jr8XQES/svyTwqmtl0GVChHAlkEmv6+VRb/7KJTCXbCWQU2lEMDOOCW/qLBGv9BQlMKMopVBx1QQwCUsJL9udUe/bYGUwp6RlEEg8g/AqPIiv9lsSL/etZTC9JCTQfGyD8Cl7iG/Nu5Iv0rjlMJYh5JBw2EPwH3OIL9/qUm/jA2VwuWCkUEwEw/AlOUfvz50Sr9JNZXCCXSQQdLlDsDGvx6/MJtKvypilcIweY9BBZkOwJ3UHb81XEu/noWVwiFyjkH5Uw7ArlMcv3ezS79LtpXCYHuNQXQqDsDNJxu/V8VLv1LnlcJhgIxBSBIOwPpcGr9kwUu/mgqWwj19i0GfAQ7A7yMZv1loS78oNZbC54qKQVvJDcDP8he/gLBLv4RdlsJvhYlBh8UNwP34Fr/2Qku/WoOWwrKOiEGApg3A8vUVv5g8S79tp5bCdpmHQX1iDcARqhS/oqNLv1DOlsKYloZBCkUNwD1JE79pZku/8vCWwuaNhUGF9wzAnMYRvwLVS7+yEZfC3I2EQequDMB8TBC//zJMv3Eyl8LvpINBUVsMwC/xDr+dyky/71qXwrC2gkHkBgzAiosNv8NeTb/yf5fC5bqBQcitC8DVIgy/SQJOv7GXl8LstoBBZCwLwDeBCr/LJE+/rr6XwjDDf0EztArA+wgJv/k2UL+s3JfCcdt9QQgyCsA8aAe/AllRv1r9l8Kn0HtBGpYJwBLpBb+E8FK/eBGYwvrLeUEZ8gjAgDUEv5iJVL/gOJjC3sJ3QfJUCMCX9QK/k0VWv2RdmMLk2XVBV6sHwKmDAb8eFli/b3qYwu0EdEFXAgfAJg8Av6XhWb9KlZjCKOZxQe8LBcDwk/2+StZbv/S0mMLUB3BBLfoEwLy4+r7L/V2/l8SYwmwgbkE2zQTAfHz3vk8JYL+j3JjCKB1sQWGpBMDRjvS+rPNhvyP8mMLQ7mlB1XwEwB478b4DDmS/rAyZwjXoZ0HzYATA15juvs/TZb9JK5nCvvRlQR87BMC2sOu+aalnv0pFmcKLzGNBwCUEwKhC6b5bWmm/+1qZwsGxYUE04gPAFDjmvlnMar8ebpnCZ65fQeqgA8B4ReO+vy9sv8ONmcLe011Bf3MDwOcc4b6NO22/ApuZwtSIW0EaRAPAuf3evlg1br/srpnCsmxZQVMaA8B8vty+Jl9vvybImcJiOVdBh/wCwPCL275G3W+/1eCZwqs4VUGX2QLAUffZvkiVcL8P7JnCDfJSQQbTAsCoM9m+kxlxv0j/UsLqGE9Cj5WMv7a+5z7fkGO/UjxSwgF9TUJkz4y/plLrPjUdZb+1alHCOOBLQnSOjb+GA+w+M9RmvxKgUMKmR0pC0caNv3xQ6z6fDme//99Pwu2sSEK8EI6/jIzrPn22Z7/HGU/CTg9HQtvgjb+eLes+YThnv9pLTsKubkVCefGNvyjA6z5Nh2e/uYhNwjjOQ0LAW42/Jk/tPh/UZr9wvEzCGyxCQh0tjb/dHO8+AgRnvyjvS8KIiUBC9eSMv8gE8T7kB2e/4RpLwlMCP0Kua4y/XGbyPqx+Zr96TErC9l49QvFqjL8OafM+Ocxmv4hyScIstztCbdeLv4tB9D6H42W/m6VIwr8bOkLgwIu/VDjzPlxlZb9qy0fCrIg4QmzEi79Q2/I+Q1Blv3j0RsK16TZCZCWLv5ID8j6LzmO/URZGwiRTNUJ8gYq/Pk/xPllPYr83MUXCUMUzQvjUib+F0vA+MNFgv2NMRMJmLzJC092Iv0I38T4aA1+/B3VDwrSpMELbCIi/AA3xPlRQXb+dh0LCXyMvQszvhr9OH/I+WnVbvyyiQcILpS1Cb8+Fv2bd8z6zvlm/qLtAwvQgLEIupIS/rCj2PjMbWL8xxz/C9aYqQrFxg78wH/k+U5lWv8zjPsI+LylC8dSBvzCT/D4TaVS//vM9wgS5J0JMTYC/flYAP8CNUr9x/jzCElQmQrbHfb8pcgI/KfJQvxIIPMJu6SRCdlp7v2qPBD8Qt0+/9Rc7whKKI0JLsHi/o8IGP5VITr9IHjrC2SUiQg9Odr/RgAg/Z99Mv14nOcLv0iBCwqVzv26UCj+QW0u/1jM4wr95H0J/T3G/8a0MP1QnSr+MLDfC3CkeQuUKb7/1Sw4/ocJIv2U4NsJ/4hxCLvBsv9kmED8Wo0e/jkM1wiaaG0K0Smu/eJARP1G7Rr/qUTTCFEIaQjN4ab+MMRM/L8FFv2BYM8JWAhlChUBnv006FT9mlES/H3AywgjPF0LtDmW/FpMWP+UVQ7+hbDHCQIgWQsbzYr8Nixg/fvdBv1pyMMKkWBVCWjZhv1j0GT+w7kC/EGMvwvggFEKSZF6/kucbPy0WP78ody7Cf/gSQv8UXb/oOR0/Z2o+v6RpLcLZyRFC9vFavx3yHj/NHD2/6n0swg6rEEIaYFm/l9IfP375O7/TiSvC0oEPQjV0V7/0vyA/Z4M6v1B8KsK5XA5C8UJVv8whIj+j/Di/TIEpwpZBDUJHSlS/yqoiP1FGOL+3lCjCXiUMQtq4Ur+daiM/uxI3vyCPJ8ICBgtCNjtRv27xIz9j2TW/R3wmwjn+CUL8DVC/XY0kP2L3NL8tnCXCz/YIQu5bTr8I7yQ/FXszv9CgJMLB4gdCWF1Nv973JD+thzK/0XsjwqbMBkJCZ0y/eSQlP/arMb+FnyLCas4FQpaiS78oJyU/pe4wv2WdIcII3gRCLLRKv2/IJD+C4C+/DpcgwrrcA0Jjt0m/jQAlP/0DL79jmB/CfdMCQubrSL+c5SQ/uTQuv0GeHsLi3AFCIQtIv5YhJT8ndS2/l5MdwjbnAEIjd0e/5FQlP+j7LL+cmxzCgsX/QaDyRr9SoCU/aZssv1R/G8Iy8P1B+VlGv883Jj+ERiy/HYQawnbs+0GyTka/ymgmP6ZPLL8EmBnCcA/6Qeb0Rb/FpiY/fxIsv/iiGMLfP/hBsjVGv3bWJj9PZCy/TpUXwjhJ9kEwXka/ygYnPwifLL/qmRbCMl30QfjERr+dpCc/rEItv6SVFcKss/JBX2NHv1LcJz+Z8i2/+p0UwuKp8EH9Ski/pvsnP7jfLr/ukBPCL8zuQTI0SL+ZIyg/G9ouv+SMEsIq9uxBUyNJv6rTJz/qoC+/m40RwrgW60EPLUq/mKAnP5iNML9/pRDCuDfpQfDNSr8Ycic/dRYxv7ylD8IwYOdB2VpLv7mHJz+DqDG/gocOwsyL5UHcQEy/X2MnPxp5Mr9Mvg3CAM/jQW4ITb9N8SY/SwszvxW+DMK6BOJBtFJOv47yJj/jTTS/YNALwgwB4EFL+U+/Sg0mP2CINb+46ArCQkfeQTikUL88VSU//N81vz7tCcJ3iNxBEAhSv9j+JD83Fje/qgMJwu6p2kHxp1K/UXIkP051N7+lcZLCDrqZQXS6QL83NjI/jqErvyxgksJNp5lBfs1DvxlGND97ey+/FlGSwnGPmUEYn0a/aUM1P/GuMr9zRpLCgoCZQZWSSL9/aDU/2bA0v7Y9ksJOeplBr4pJv1IuNT8CkTW/yzSSwm5xmUGuAkq/yZg0P+3KNb/kK5LC6GeZQeGOSb+QIzQ/aSY1v8ElksI6bJlB/A5Jv5HZMz8TiDS/LSGSwpFymUHn5Ee/TaAzP4FHM7+rG5LCrHWZQZMjR7/4jDM/WX8yvwgWksICcJlBrHxFvwLOMz8i9jC/VhOSwll8mUFN1UO/RfQzP0piL7/dDpLCbn+ZQVTcQb+wvzQ/bcAtv9kMksIzhJlBQMRAv6MYNT8azyy/mQuSwouOmUHpTj+/F9c1P6WpK7+eCJLCdYOZQeUfPr8wQTY/Rqgqv8QHksJmjZlBKcI8vw45Nz+jrym/3QKSwvaYmUGSDTy/XBs4P8JUKb9qApLCUJmZQbHVOr9ydzk/PqYov2ICksL1mZlBmNU5v1i0Oj++ISi/2P+RwpKemUEGlzi/Azw8P4J6J79a/5HCPKuZQf39N792jD0/UWEnvyAAksLuo5lBdr02v2x3Pz9j2ia/Wv2Rwl6mmUHb8DW/bh9BPy6sJr8B/5HC2reZQRdONL9R3EI/O64lv2IAksKrtJlB4Ugzv8F9RD/LQSW/4v+Rwp3AmUHQ3jG/kxZGP4ZrJL9i/ZHCdrqZQdO4ML/nYEc/nbsjv2ABksLLwJlBpI4vv4i1SD/3CSO/bwSSwo3EmUFX7i2/1tlJPxLPIb8RA5LCBMCZQVBVLL/UAUs/pJsgv1kFksLLzZlBWeMqv+TaSz/9ch+/4gSSwtbUmUGq2yi/76dMPwWvHb/sB5LCGNKZQUpQJ78eTk0/WFocv9EJksJ+0JlB8wMmv7TXTT8LOxu/gAmSwgXRmUEu/iO/9ElOP9JZGb/KDZLCR9GZQbUtIr9V0k4/i7UXvygQksJ03JlBI0AgvyJxTz9x+xW/TRGSwubemUGHmh6/xwVQP92FFL8hFZLCVuSZQUifHL9cjVA/f7YSvzMXksKv6ZlBTqUavzgnUT8s7hC/jhaSwjnimUFu1xi/vM5RP99VD7/mF5LCGuiZQRNKF7+emlI/kAgOvycdksIY6plBAJMVvxloUz+VkQy/jiGSwjLumUH/2xO/NfFTP70FC79QI5LCAeyZQc6wEr/MGlQ/negJv8ElksJ07ZlBprgRv2GnVD9OGwm/niiSwhz2mUEhGRC/Z7ZUPyKDB78bJZLCO+6ZQZunD7+H6lQ/kiEHvzIoksIM85lBC5gOv2uYVD/j/AW/Ny2Swt3wmUG11A2/2qJUP4s+Bb+uKpLCyPCZQe63Db8KrlQ/QSUFvxsrksJ15ZlByI4NvypuVD9t6gS/bimSwiremUFDeQ2/Lj1UP0PHBL9NKpLCwNqZQRyVDb8/u1M/6L0Ev4cwksJH35lBGEQOvwhIUz/WSQW/3iiSwm7fmUHVlg6/0AFTP26HBb8wKpLCSNqZQeRcD79Iu1I/qDYGv+gvksIp65lBktcPvzmYUj+npQa/IDGSwj3dmUF9vxC/WGZSPzt8B78uK5LCc+KZQaxLEb8DzFI/ZiQIvyItksIq05lB93MSvzREUj/DIQm/YiySwv7FmUGgXBO/crNSPwUpCr9UJpLC18uZQV+JFL8YsFI/Y1ILv04sksJO0JlBcq0VvzDmUj/QhAy/JiuSwoTImUHg0Ba/BytTP7m7Db+8KJLCctmZQRmuF7/y5lM/oNEOv5QpksKPxplByvEYv/TFUz8VCxC/1yaSwhK1mUGR7xm/mwlUPw0eEb+eIZLCfrqZQQP4Gr/kH1Q/KC4SvywfksL/rZlBKigcvxAuVD8LZBO/sh6SwjasmUEaSB2/gTJUPwuHFL84G5LCHrOZQeMrHr/U/lM/8FsVvwgcksIioplBUkwfv/t/Uz/ZVRa/fBySwjaSmUECcSC/+hJTP59ZF7/SHpLCC5yZQfucIb/lAlI/YS8Yv5UYksIneJlBVtIivyxMUT8RKxm/WRqSwmJWmUGp0CO/9WdQP8rfGb9DHJLC4FeZQfIDJb/EKk8/basav+QZksLlMZlBWBcmv+0yTj8xbRu/v3RWwvKXVEKTC32/DucJP10/VL86i1XCyBBTQkVWfb9mvws/ZohVv7aQVMLQh1FCDeR+vxAfDD8FSle/vJ1TwpUCUEK4e3+/UtgLP5C7V7/rtVLCBHtOQuAogL+MAww/CKpYv93IUcK670xC5AuAv6vvCz/rZFi/uNNQwkphS0LOJYC/GFkMP6fSWL8o6k/CA9NJQhVCf784RA0/5UdYv2j3TsJjQkhC9Al/v6xUDj+ho1i/5ANOwiuyRkI5in6//HQPP16/WL+vC03CJDpFQkmefb+7VRA/wUpYv/wYTMIYqUNCNbJ9v4L4ED/ctli/YhtLwvETQkJqrny/B3kRP2L1V78aK0rCyopAQimbfL/T9xA/cZxXvxgtScJaCj9CLcJ8v6i7ED92o1e/qjRIwvt9PUJUk3u/Xj0QP5cuVr/bM0fCWfo7Qrx2er80vA8/+8tUv58sRsJCfzpCCjp5v75PDz+pVVO/PydFwmz9OEKUhHe/+kkPP6meUb8gLUTCeIo3Qj8Kdr/+9Q4/FftPv1kfQ8KNFzZC3SR0v5Q+Dz+AQE6/GBtCwryrNEKxIHK/9dAPP4KOTL89FEHCiDszQsQMcL8fthA/kPdKv3H/P8IE1DFCjOptvyzoET/xeEm/XP0+wrJuMEL9BGu/CVITP3RVR79r7z3CIgovQgg/aL+nCxU/u3ZFv1PcPMJnuC1CHLVlv1HVFj852EO/W8k7wkBfLEISmWO/wpkYPzigQr/bvTrCbBErQr9VYb9QXxo/xT9Bv/SnOcIivylCBDZfv5C2Gz9DzD+/lZc4wgJ9KELa51y/8lwdP15OPr+siTfCVjQnQoPwWr/OAx8/ciM9v4ZpNsIJ9SVCq/5YvwAuID92wzu/HF41wo2+JEIbOle/5ZYhP7uqOr/QUDTCY4UjQhf5Vb+AhiI/Vts5vwxIM8LUPCJCkYdUv9e4Iz8r+Ti/QDkywp0MIUI8p1K/ZmAlP3/cN784OzHCPegfQmbMUL/AVCY/m3U2v80jMMIvsB5CkxhPv4XmJz+ddzW/yxUvwhCPHUJFqU2/qPcoPzuENL+58S3Cj2YcQpVES7/yiyo/oNYyv9r1LMLIShtCW1ZKv5qEKz8sVTK/UtUrwpgqGkKXkki/euMsP+QrMb+82CrCYRgZQotcR78Ccy0/Czcwv1nTKcK6+xdCUMtFv7IQLj957i6/dbgowg7kFkLAC0S/GB4vP3OlLb8uqyfCw9QVQsBwQ7/iVC8/1CMtv96xJsINxhRCEkxCv521Lz9eLCy/+Jwlwr6zE0KxL0G/L/YvP0YwK7+2fSTCdrYSQrNmQL/1OjA/Uocqv9WQI8KWuxFCvhg/v4lOMD+BSSm/yogiwlOyEEK1gD6/1gkwP7SaKL+SVyHCzKcPQpbyPb8t4i8/JQEovy5xIMIQtA5CAI09vziJLz/+eye/02UfwtrNDUIf/Dy/S98uP9GtJr9eVx7CpdcMQsBpPL+TyC4/gRcmv0BQHcJ22AtCiQU8v3FdLj94jSW/jk4cwlLqCkIOkTu/rVAuPzcYJb9sOhvCeAAKQoJPO7/XOS4/SNAkv6E9GsJTBglCjis7v949Lj8uryS/WB0ZwkgnCELf+Dq/P4kuPy2bJL99HBjCRi4HQrc5O7/ldy4/C9Mkv2wsF8JISQZCOUk7v4xoLj8d3CS/rDQWwpprBUJkzju/fmMuP5haJb/GJBXCiHkEQs9aPL+VSC4/vNclv5QlFMJijANCZQ49v+SwLj+KrSa/gh4TwtXBAkKT8j2/RbYuP3iMJ7+FJBLCqcQBQis7P7/qiy4/cLoovycZEcLR3gBCxnU/v35+Lj8I7ii/ThUQwrr9/0FknkC/AvItP5jWKb+EFw/CjS7+QRrtQb9tiy0/BvMqvxgxDsLTZvxBocRCv0kzLT9goSu/lTENwsSm+kHGi0O/IA8tP5pULL+GFgzCf+D4QUuvRL+QwSw/NVEtvxtPC8IYNvdB67BFv3YsLD/JDy6/kk8KwtyA9UH3Hke/JwwsP6hnL79cZQnCeo/zQUbqSL/lFSs/3MIwv0h/CMK85/FB/bVJv0BMKj80NjG/pIYHwh478EGCPEu/WN4pP/GFMr9hngbCZ27uQWz4S7/dTSk/AwEzv1aKk8Ip76BBslUTwFu0Lb8Ywz+/+vOTwpogn0HMsBLAwQgsvwObQb9UXZTC7WSdQeF/EsCp1Cu/iUhCv1bHlMIVuZtBJY8SwCtmLL8ATUK/UC+VwgoYmkE36xLAca0svwb6QL+WlJXC8nSYQZh4E8C1my2/tys/v1bylcLz1pZBsfMTwF84Lr+Ugj2/BlaWwuhBlUGMihTACdYuv/ZqO78AtZbCMq2TQaHkFMAX5C+/Znc6v/EWl8J8GpJBmUEVwBKfML/KUzm/TG2XwjTBkEHinhXA31Qxv14sOL+IyZfCbTKPQWHJFcCvGDK/hdU3v7semMLcnY1B/QIWwFwAMr8I5Ta/bXuYwqQWjEH7GRbA0n0yvwe+Nr9r05jCE5OKQRgJFsBstDG/raw2vwEpmcLS+ohByv8VwDMiMb8rlDa/sXmZwnSGh0EfGBbAeBIwv+vANb/szZnCxwyGQdgBFsBWni6/nXw1v9wimsK4doRBT+oVwIaOLb8aZzW/rXaawlf+gkGuzBXAzrQsvyGANb/SyprC6nOBQW2yFcCy6Su/q5E1v/oYm8LL+H9B23wVwEuFK79dOja/C2ObwvADfUHGYBXAQHMrv4ehNr8etJvCTvp5QZ84FcAYsyu/ils3v3sGnMI6LXdB1kIVwPwCLL9eVTe/Z1GcwoJQdEEnLhXAEmosv3nTN7/IpZzCZ3BxQSYcFcDnFS2/sGQ4v1bnnMIMnW5BEe8UwCtrLb+cPDm/7jidwi7Da0HwvhTAt5ktv2UQOr9jgJ3CZPRoQbyiFMDg9S2/3ag6v/TNncIfImZBKmMUwITQLb+Olju/8w2ewhBdY0HtHRTAwEEtv+xsPL+7VZ7CkoBgQTHnE8CLGi2/tzY9v0qonsKOwV1BjpkTwIACLb9BYz6/fvGewqwYW0HFVBPAsJosvx5JP79rMZ/CaE9YQWwcE8B6Iyy/HPY/v959n8JzdlVBi9QSwMh8K7/2y0C/rs2fwinHUkFtrxLANP4qv/YnQb+vFKDC8/9PQZttEsDAZyq/SOxBv2ZXoMIyR01BkDASwHvtKb8CqkK/spigwr92SkFRFRLAdzwpv5TGQr8v3qDCI8RHQdLUEcD2vCi/VI9Dvx0aocKYCUVBypsRwHSjJ7+q8kO/9WOhwuZnQkFSeRHAS9Ymvw8eRL/ErqHCMb0/QVBpEcD/eia/7zNEv1zmocLVAz1BI1wRwHmUJb+k/UO/FiiiwtRhOkFrIxHAsKokv9JzRL+GaqLCIKM3QYkeEcDd8SO/CDFEv3ylosKu4DRBXvgQwD0fI7/2ZkS/mOKiwgFGMkEGqhDA8fshv9IWRb+QHaPCK4ovQRyFEMAZuiC/ChJFv3tVo8K1qyxB+C4QwIpCH7/StkW/ao2jwuDiKUGb2A/Ax+Mdv9BmRr+aw6PCQFcnQbdxD8CffRy/glNHvzICpMJ2qyRBgxMPwNg1G7/yKki/fjykwhPqIUFetA7ALNAZvyz2SL+LZ6TCfwMfQRkrDsD+NBi/tExKvz6kpMIAlBxBjqoNwFm0Fr/Pi0u/XtakwrvTGUFJGw3AhyAVv/35TL9KCqXCb+wWQVF9DMAquRO/gbdOv4gxpcJ8DhRBSdYLwLj+Eb/KbFC/e2+lwtIgEUH+NAvAXNoQv8RXUr8MqaXCxFsOQVSKCsDHgA+/oEtUv8vYpcKAqAtBt+EJwOIFDr+AJFa/bwemwl6ZCEGlNwnAEugMv6Y0WL/fPabCidgFQWF6CMBngwu/HGtavwxfpsLUFQNBJr4HwMkKCr+5kVy/aommwkUnAEG5EwfAGJUIvwdxXr+lvabCZBb6QCtWBsBCBge/jY5gv73gpsJwAvRAp7kFwOCuBb+8RGK/RROnwvp57kB+IAXAYGgEv0z2Y7/uP6fCHDHoQOWHBMBeOQO/1LJlvy1op8KeH+JAbvYDwKieAb80EWe/Ro6nwoxj3EBAcwPAiiEAvyVFaL+AwKfCOtLWQIdwCMDfUP6+bSRpv8rhp8KkbNBAwz0IwP9Y/L7XEGq/dAWowqFlykBO/AfAPPP5vlQea79ZNajCkjnEQH/SB8AO4vi+hV5rv8peqMIYer5AmJkHwOX39r5IHmy/DHuowpkcuEDfkAfAgF32vg+CbL98AZnCSWMGQZPRFMCJpzK/oPU7v++3mcJXOAFBr4sUwFWIMr8GAj2/vG+awgRy+ECpkhTAVU8zv2A8Pb//J5vCwZXuQPCrFMCpZTS/7049vyDim8IByeRAB/8UwA3wNL9UOjy/X5acwrPe2kBJiBXAjO41vzZ8Or97Rp3CmRfRQAEOFsCWfTa/6Jw4v3T8ncKUW8dAC7UWwK4gN79dQDa/YK+ewniHvUAfFhfAQEQ4v28yNb+iY5/CBO2zQE+PF8D3Bjm/CpszvyYLoML7rKtARCMYwHEHOr/ksDG/6rugwsRBokAWiBjA/j87v3uZML+WYKHCbJeYQF8HGcCEqzu/4MUuv+YSosJQVI9A+loZwNzJPL9q5y2/YL6iwpP2hUAYrRnA7bM8vz6WLL9hY6PC7sR4QN7uGcD+5Dy/c6Irv84DpMK54mZA6l0awMiPPL8Uxym/D6qkwtDQVEAZqhrAEq47v45CKL+cUqXCGQ9CQLDfGsDANzu/kkEnv+z2pcLJAjBAQhYbwBv8Or9yUya//J+mwprDHUBwPxvACtA6v22gJb+sP6fC0RYMQBE8G8Dq0Dq/Dq4lv7bcp8LSKPU/XmobwKobO79mEyW/F4Kowkhf0D/PYhvAMdU7v1t2Jb/iKKnCLA2vPzuaG8BO2Ty/F/wkv//IqcL8Qo0//L4bwCb0Pb9h0yS/tXKqwp7pVD+hxhvAsHE/vwdCJb8ICavCz3QSP2G+G8APvUC/Sd0lv0Kwq8KuPZ8+trUbwI74Qb+KdCa/QkuswmgVYz0/zxvAMHVDv/2aJr/Y8azCsrZRvjy9G8CKekS/y0Mnv92IrcJhzuy+o6gbwJguRb9E2Se/uCauwlDNN7+DqxvAn3xGv19JKL84zq7CVF14v4OcG8BXsEe/C/gov7p1r8KSi5u/EJ0bwOWbSL8pTSm/BgWwwh7Iur/DqxvArXhJv0RjKb9yr7DCV4Lbv+mmG8CjL0q/wropv2hVscIbivq/x84bwM7gSr/tWSm/wvexwrLpDMAA1xvAGqNLvxiAKb9jkLLCzk0cwGTzG8AxU0y/Fk0pv1gvs8K22SvAZSccwBrQTL+bpii/SsuzwuR+O8ADQRzAJ2JNv2BzKL8fX7TCOkdKwBxvHMABqE2/jdAnv10BtcID4VjAnaocwG4BTr8R/ia/r6S1wmx4Z8Cs/hzAq/NOvzT+Jb+wMbbC9YF2wKg8HcD8YU+/ryglvyrKtsIcfYLAV28dwKTAT7+SeyS/CGu3wt7+icC3xB3Alz5Qv/VLI7/9/rfC+cqRwEH7HcCmzFC/gp8ivzicuMIsvZjAcRoewOD/UL9QMiK/Siq5wjSmn8DUUh7AZPRQv7VIIb+SuLnCE2inwIlnHsA40VC/SOggv1VOusKw4a7AKXAewJDQUL/5xCC/XN+6wqNztcAZXR7AX2hQv2TuIL/ocbvCEtC8wAVhHsCBwVC/Vv0gvw0JvMIaH8TAK2EewJl0UL8X4iC/Z4q8wvXVy8CDPB7AxzRQv+9gIb+mGr3CM3jSwNEaHsATyk+/vMQhv3CqvcL+z9nANtsdwBqDT79/riK/DjK+wuRN4cBmmR3A3k5Pv+unI78Gr77CX8zowC1ZHcAf0E6/gYAkv+JDv8JzdPDAHQsdwN+9Tr8muCW/fNe/wp3h98B6rxzAmoJOv+MYJ78gWMDCUtj+wKBtHMA0+k2/fvQnv3fgwMJydQPBFhYcwFr1Tb/FWCm/lm7BwmMKB8HkqhvAkG1Nv+LdKr945sHCH5sKwfRBG8Cz/Ey/WWIsv2RjwsI4XQ7B+/IawAYXTL9RUC2/oObCwr4wEsGDiRrAFTJLv6aqLr9bYsPCnU0WwWtEGsCijUq/qocvv2Ttw8LKvxnBowIawEHdSb8jUjC/NWvEwvy0HcEFrhnA3ABJv4lYMb9T4sTCzW4hwQRvGcBMn0e/jdExvypjxcLIASXBJ0kZwNiLRr8RATK//OHFwkKdKMFQQhnABNJFv0LUMb+QV8bCflQswY8iGcCJJEW/BBIyvzTIxsJkCjDBPP4YwBzkQ78yKDK/IlDHwnC3M8FuHxnADI5DvwJ/Mb8Ky8fCPkE3wX//GMCYhUK/4Jgxvxw0yMIWDzvBRRgZwPNoQr/lKDG/e0OawpZm9kCR/RTAXCUzv516O7+v+5rCnf3rQKG1FMB+ATO/TY08v4G1m8Lp7eFAgbcUwAjBM7+92Dy/92+cwk4A2EBUyhTAsc40v1IBPb9sLJ3CxCDOQBYYFcD9UjW/ev87vwXjncJZI8RAr50VwDpONr9rTjq/ypWewghKukCyIBbASt42vzF6OL9eTp/C4nuwQADGFsDkgze/TyU2vx8EoMKblKZACScXwL6oOL+4FzW/MLugwn7qnED8oBfA/285v79+M7+KZaHCW6CUQLA4GMADeDq/I4gxv1AZosLNKItAoaIYwNq6O7/ZXzC/AsGiwipwgUACKBnAnDM8v7R4Lr9+dqPCFEBwQNiBGcBOYT2/x4Ytvz4lpMI1Zl1AwNwZwMhgPb+lGiy/Ms2kwjEnSkCzJhrAY6M9v7wMK7/CcKXCMi84QMudGsC9YT2/wBgpvwcapsKdAyZAxPEawLKPPL+meye/1cWmwqcvE0BJLhvALiY8v0NkJr8SbafCPA4BQDdrG8Bn9zu/6WElv1wZqMKkft0/p5gbwAbZO79DoyS//LuowlcHuj9KmRvAv+E7v/6jJL9gXKnCpN6WPzbLG8CWNTy/G/4jv7UEqsJw3GM/0sUbwOn3PL8nWyS/qK6qwvIIIT8N/hvAzwc+vwzhI7/FUavCHJ+6PgQkHMBDLT+/Y7Yjv4f+q8KHsbs9KyscwFG0QL8eKSS/AJiswlF/LL6nIhzA9AlCvxDIJL9LQq3CshXcvroZHMAQT0O/zGIlvzDgrcKady+/BTQcwFzWRL+1iCW/JYquwoEmcr89IRzAYONFv5A2Jr83JK/CYh2av8AMHMDZo0a/l88mv1PFr8ICxbq/RBAcwIH9R782QCe/km+wwsUI279aAhzARDpJv8vsJ7+IGrHCK2L6v0cEHMDJL0q/JD8ov8CsscLMvAzAdRYcwJcaS79hSyi/tFqywgwSHcD0EhzA/95Lv6ihKL+4A7PCoIcswGg+HMA4nEy/vDUov4yps8J9FjzAYUocwIFsTb/kUCi/aUW0whhkS8DgbBzAjitOv4EJKL8r6LTCzNNawCemHMA3t06/AVInv06HtcKuZGrA68YcwDVUT7+0BCe/pB62wsYLecBB/RzAybBPv81HJr89xLbCQ7yDwFRBHcCDGFC/plYlv5dqt8K28YrAw58dwNodUb9xMSS/kvu3wnZhksDo5h3ARKJRv1o9I7+Ml7jCWYSZwCwmHsAuGFK/LWQivx88ucIy7KDAMIgewJWsUr+RByG/ONS5wribqMDIyx7AzlRTv6QtIL9wdbrCh26vwLf6HsB0pFO/aokfv0cHu8IAMLbAH0IfwLqyU787ax6/YJm7wn3OvcDPZx/AoK1TvxHQHb9bM7zC8iPFwIiBH8DHx1O/NHAdv2DIvMK2h8vAvX8fwNx6U7+aXR2/b169wg690sB7lB/ADvhTv1kzHb8R+r3CWt3ZwF+mH8BPx1O/ONocv4V/vsIGY+HASpMfwEemU7+4HB2/RRO/wvbT58AUhB/A5FpTvzNBHb/zp7/CNPnuwI9VH8DPNFO/gfEdv8szwMK1OvbA6iUfwGQiU78hrR6/7LTAwkuA/cDO9x7AOslSv3lKH7+ATcHCPXYCwdW6HsCA1lK/cUcgv5DlwcJXDQbBs28ewJ+9Ur9ncSG/8GnCwl1nCcGYPx7ADFpSv0sTIr9k9sLC7U4Nwef3HcCqdFK/k0Ejv8eIw8KPwBDBj5wdwHkPUr/dkyS/5QTEwgMsFMFLQh3AE71Rv4ToJb/OhcTC48cXwZgCHcDQ9VC/eKYmv1EMxcJEchvBF6ccwAkrUL+e1Ce/2ozFwkppH8HubhzAyKJPv4OJKL+kG8bCubIiwds4HMBPCE+/8y4pv4CdxsLGeybB+O8bwOk9Tr+YDyq/yhfHwhILKsHIuxvAh+5Mv7JpKr/TnMfC3HItwbygG8BO7Uu/GXkqv1QfyMLW4zDBMKMbwHU7S78qLSq/XpjIwuBpNMHgiRvAAphKv8RXKr+6DMnCavE3weJtG8AUXkm/BFUqvzyYycLgcDvBaJMbwHgLSb/AnSm/2hbKwtzMPsHBeRvA+gdIv3elKb8wg8rCeWlCwcuVG8Aq7Ue/0ykpv+GKmcLwhRpB4+ArQFvdab+wU+M+ph6awkKsHEGlvCtAfK5qv6XZ5D4DsZrCRcQeQanKK0CjtWq/rGvkPi9Dm8I0zyBBnLcrQER/ar9P7OQ+zNSbwgPvIkGceytAoKZpv3Nr5j5iYpzC0P4kQVhpK0Aa8mi/yanmPjDrnMLSECdBjVIrQB34Z79Z6+Y+IXedwktCKUFGLitAgJZmv1hn5z5qAp7CoHArQfUeK0BnF2W/zivnPtiMnsLIoC1Bkh0rQKnUY78aneY+vBGfwj0NMEEtHStAM89iv58j5j4Ll5/CukMyQYtBK0CKc2K/RNbkPnMaoMJaYDRBOFMrQGycYb/A4+M+DKCgwjiENkHdjitAa7Fhv9sT4j4pK6HCdpQ4QYLWK0CPHmK/8g3gPjqqocIBmTpB7/4rQOSgYr9KCd8+Lyqiwq2fPEEPOCxAq0VjvzmP3T7Dp6LCB48+QZNzLEAmqGO/MuPbPvElo8I2cEBB9rIsQBgXZL+fHdo+oJ+jwmJJQkHH5ixA7tJkv/DV2D7AG6TCoiREQZgfLUB+emW/iFzXPqqJpMJV6kVBzlYtQIDpZb8d1tU+egelwj+UR0HBlS1ANE9mv50N1D7pdKXCTiZJQbzeLUBTz2a/QwDSPpbnpcKq5EpBFjwuQI0nZ78jP88+Y0qmwomFTEENgi5A7x5nvxwPzT6JwKbC8gdOQa7NLkC/Vme/gM3KPnIup8KPeE9BdRUvQM59Z7+co8g+MJWnwj4QUUF0Yi9AJHZnv3Q9xj6D+6fCcKtSQdyZL0Di7me/1bbEPstfqMJVFlRBnMEvQPLdZ7/DdMM+g8WowsuRVUHp3i9Ati1ov0mswj4bK6nCSx5XQYj4L0DVgmi/6QLCPsSNqcJwh1hBXvovQBPcaL/HF8I+ZO+pwtbrWUGB0i9A5c9ov39Pwz6qSqrCP4ZbQR6vL0CRSGm/GZnEPh2sqsKM5VxBk3gvQFk8ab8YRsY+ywWrwqZ4XkG4HC9AsBlpv0YTyT5gW6vCsghgQRujLkDk/mi/etHMPgG3q8IIk2FB5ykuQGWmaL8kc9A+zAyswlYlY0HEhC1AIwtov+xW1T72W6zCpKRkQa/FLEAFPGe/jfLaPsCsrMLMA2ZBBS0sQNj8Zb8tJt8+sP6swqa/Z0HXZCtAjS5lv0EH5T50Tq3C2lZpQaKAKkBu3GO/HIjrPtmWrcKo82pB0rEpQI6FYr9iV/E+auWtwrmMbEG50ihAUkhhv2qy9z70I67C0R1uQQrYJ0AS51+/0NT+Po10rsLKzm9BMwEnQJL6Xr++iQI/dbSuwjqXcUHBDiZAdbddv7X/BT/F/q7CWFxzQfspJUBIpVy/j0sJPx81r8LAQHVB3S8kQFJpW7/W4Aw/RXivwpAQd0EZRCNARgRav9guED8qva/CEAt5Qb1sIkAq9li/0EMTP/f3r8Ix5npBt44hQGPIV79yaRY/XD6wwr7hfEHuvCBAXD5Wv5o+GT9rcbDCbuB+QWf1H0BzGVW/qwkcP8KgsMIIfYBBxBkfQIRIU7+O6x4/Ee2wwgSOgUG1Zh5A2+VRv2lLIT+KJbHClqiCQRKeHUD2IVC/bN8jP/BascJrsINBcescQO+tTr+qMyY/bpWxwqK0hEGVLRxAspJMv+F2KD/g07HCtuCFQcObG0B4Bku/0TcqP4QQssLFAodBhREbQLRXSb/Tyis/wkOywrQoiEGBjxpAQ4xHvwkvLT8thLLCmE2JQfUSGkDQf0W/b2EuP7G4ssLDhYpBc50ZQMiUQ78mgS8/1vKywsOVi0HvQBlApw5Cv11gMD8VGbPCNLmMQVPfGEB13j+/Rw8xP7Vzs8KQ3Y1BKrkYQAhsPr/MFjE/AKCzwqIaj0FTchhAZ7s8v0iIMT8H17PCElSQQac6GEBtujq/5poxP978s8LBdZFBkiIYQMEIOb9ITTE/lzW0wtKdkkGL9hdAF6A3vwpsMT+Va7TCOtaTQQLmF0DCgTa/XjoxP9SatMKv7pRBivkXQJzHNb9woTA/rta0ws/6lUG2JRhAK3I0v6ZoLz8OC7XCAPaWQSFUGEBXwjO/i2ouP41FtcJfFphBcnkYQMYcM78rlS0/rWm1wjIwmUGYrxhAzUkyvy9sLD8CpJfCsw9hwVwBBED3dwC/BTRmP5zgl8ILf17BgYcEQMS8A79gAmY/ZRqYwvf6W8GDYQVA2Q0Gv5njYz/CUpjCMoJZwRMOBkDgMgi/T2FiP3SJmMIKBVfBRLYGQN7XCb9spGA/D72YwkSkVMH0bQdA7ocLv/msXj9+75jClVFSwd8uCEACPg2/uJFcP74imcLS70/Bb64IQMV0Dr/gN1s/WleZwoybTcE0UQlAc2QPv+snWT+0iZnCz1VLwVDrCUBIjRC/FlpXP0+8mcKo2UjBs28KQHicEb+Z1VU/PeuZwg2KRsEpCQtAzf4Sv60nVD/hG5rCEV1EwfV+C0BwlRO/g51SP2VMmsJ+OkLBViUMQFPLFL+Jo1A/AYaawr0cQMGIrQxAYwcWv9AkTz9xsprCfR4+wWVBDUDwHBe/EWNNP5jemsIgEjzBILsNQKX6F78U7Us/KAybwgYjOsFbPQ5A8HkYv08mSj/YOpvCbD84wVapDkB03hi/WatIPyRlm8IBejbBBhAPQMx2Gb8gX0c/WJCbwjyyNMFKcA9AkPcZv/EgRj91s5vCpg0zwd3VD0CLZBq/fcREP57fm8LLfjHBZzgQQHrWGr+xdkM/tAKcwov1L8ESkRBA6FYbv85WQj8xKZzCAV4uweoNEUC1/Bu/fbpAPzJDnMLY6SzB9EYRQMX2G78k2D8/oWucwlmQK8HBlBFA/kMcv4PLPj/wkZzCIU0qwZ2+EUBxMhy/0R8+P52wnMLK3yjB/OMRQL7+G78rdj0/z86cwuZ3J8EO9BFAqggcvxQ8PT9E6pzChDYmwR/nEUAmZBu/3iI9P0IKncLX4yTBAOsRQGBHG7+JBj0/AimdwiSmI8G3xhFAJ6MavyNIPT/gQJ3CaGsiwbWcEUBVahq/WdE9PxZgncLFLyHBFFgRQPDsGb9qoj4/Un2dwhbaH8FvExFAubwZv4SXPz+vmZ3CJa0ewWKpEEDIBhm/DeBAP+u3ncLZUB3BHFMQQICSGL9B+kE/CcudwrsIHMF6yQ9APFkYvzD6Qz/b5J3C8c0awcJfD0Ak2xe/xVxFP637ncLphxnBQt4OQAtwF7+pJkc/wBCewqZGGMERTg5Ak8EWv0oKST9iKp7COUcXweznDUCW9xW/LjpKPy1CnsI8+RXBy2ANQC6LFb+lG0w//FeewmS/FMGA1QxADNgUv9jqTT/baJ7CBIITwaVfDEBN+xO/nE9PPz6CnsJYYxLBEewLQP9RE7+mxVA/Q5Gewt5CEcFObwtAuOESv2x+Uj9dqJ7C9xwQwUQaC0BmbBK/75VTP2a4nsL05A7BsbAKQBv9Eb9bA1U/bdCewqnWDcHuYgpAargRv+wXVj+52Z7CLY0Mwf7iCUD4JhG/VM5XPyTwnsL7hgvBBKwJQDjAEL8xdVg/UQafwlNWCsH5aQlAv6QQv6pxWT9wFZ/CJFMJwSc2CUA1iRC/1DRaP2Qxn8JqFAjBfvwIQMVfEL9dCFs/c0Kfwu7/BsGOzAhA5BYQv1OjWz8PUZ/CdOMFwfuKCECoog+/Pm5cP1pwn8LPrQTBbHoIQBmXD79Qq1w/WYmfwouDA8E5RghA3kIPv3NRXT/qm5/CVKkCwdgYCEADEQ+/5+5dPzCzn8JrgAHBHO8HQIh9Dr8fR14//cyfwl1uAME73gdAE1kOv653Xj/n6p/C4cX+wAnTB0ASWg6/9KVeP2gKoMLwavzAXMwHQCgTDr+TmV4/jiegwjY8+sBGugdAdckNvyu6Xj8kQ6DCGff3wFjEB0BmYQ2/1VZePxVeoMLwFfbAUN8HQHR1Db8o9F0/qW+gwp3m88DJ3QdAlL0Mv6+TXT8ro6DCoMvxwC3+B0BO1wy/YR5dP5C1oMLKMu/AtggIQPB/DL/vwlw/HNWgwtwx7cCrKghARv4LvzXxWz/46KDC4gDrwOQ7CEC2Mgu/lTpbPyMIocIPxujAZlMIQDzVCr/vp1o/HSKhwgxx5sCvbAhAGHwKv8kQWj9lOaHCfHvkwNCtCEDbMgq/u+JYP2RcocL2suLAXN0IQBGnCb9511c/In6hwsMV4cCDCwlANCUJvw3YVj/rlaHCKzLfwOswCUCsrAi/PgFWPwKsocICoNzA01AJQG7ZB79/D1U/w7WawizGL0FvWhhAeR89v5IQMj808prCCIcxQWwBGED4sT2/G7MzP3Uvm8IKODNBHP4XQK9dPb+FnjM/Xm2bwnHdNEGp9xdAAtg8v6uCMz+ZqJvCf442QQnvF0Dj/Du/CU0zP9bim8J7LzhB0hMYQMBmO79GfDI/SBmcwkLKOUHKSBhAvfk6v4R7MT+xUpzCKns7QQltGEDpUTq/+qYwPz+MnMLvJD1Bu7IYQHTdOb+7YS8/eMacwrbIPkHRDxlAxtc5v/TrLT8Y/pzCXp5AQQRrGUAQ/Tm/Uo8sP1YzncJ6RUJBk+8ZQA28Or9Tyio/GGmdwuzYQ0F4YhpAzu06v2IVKT88oZ3C5WRFQb/+GkDH5Du/3QcnPzLgncKq70ZBj5gbQO4HPb/9EyU/2hOewtRjSEHTIhxAzww+v31RIz84SJ7Cyd5JQU2oHEDM3j6/tY4hP3p9nsK6SUtB9CgdQKB2P7/JyR8/wLKewiSlTEHpnh1Ahtc/v6cbHj845p7CXv5NQSwDHkA4fEC/IsscP54bn8LeT09BuW8eQIX4QL/LSxs/hkifwmSPUEE10R5ASmNBv+vxGT+WfZ/C7rNRQUM0H0D4wkG/Mo4YPzGrn8LA0FJBypkfQHpUQr+vMRc/x9mfwj4GVEETDiBA0eZCv42bFT8R/p/C3ihVQQRXIECR5UK/G34UP6sxoMLRMFZB5bAgQJkdQ7/BMRM/S1+gwqgoV0GJ8yBAEDdDv1I2Ej/7iqDCBUlYQREuIUA2GkO/KEkRP061oMKibFlBN14hQEpOQ7+nnhA/LdugwkRyWkGJXSFApuNCvx9/ED9uBaHCqoNbQXlnIUCu0UK/xVIQP0IxocJ8nVxBblchQMikQr+rghA/zFahwq6kXUEuJyFAmK1Cv/FAET9RgaHCka1eQVHiIEBZU0K/ei8SP/6mocLG319BdpIgQAt0Qr/6cBM/O9Ghwq7uYEHhHSBAoxdCvx4ZFT/w+aHC0ShiQR+fH0BQu0G/aukWPwUdosJAYWNBpugeQF6XQb/Kpxk/wESiwviVZEEYTx5AOQ1BvyzTGz8kaaLCIORlQSR/HUBKWUC/psYeP2iMosJ6GWdBEaIcQL18P7904CE/jrCiwugkaEGH5htAEBI+v99BJD9l1KLC1YZpQZUHG0CkID2/o1snP//4osIn2WpBVRcaQLLAO7+bkCo/WBmjwtYnbEGwOxlAhjM6vx1hLT97P6PCUG5tQd5jGEDg3Ti/BDcwPzpYo8JjuG5Bln4XQEWON7+ZRDM/wHyjwjIVcEEGwRZARoM2v+/ONT+Dm6PC0oRxQTL3FUCNTDW/PHg4P7a+o8KA43JBeT8VQOwxNL9F5Do//NOjwsZ4dEGqdBRA3BUzv6mcPT8t96PCF9N1QXfQE0De6zG/mbI/PxcYpMIUXHdBDToTQKQXMb9atkE/JzOkwjLCeEHftRJAnE8wv5Z1Qz9iWaTCEEx6QYE6EkDlWy+//fxEP5BzpMIg2XtB6M0RQC2WLr/eXEY/lYqkwgd9fUEjSRFASG0tvxHxRz9xtqTCthd/QfEAEUA80Cy/jM5IP1bZpMIzYoBB+JoQQJjLK7+j9Ek/DvKkwnQTgUE/RhBAWAIrv6/vSj92FKXCvt+BQYL0D0As1Cm/0K1LPzo4pcK9sYJBJ8IPQP4FKb/OGEw/aV6lwt2Ag0ELnA9AQE4ovyBcTD9Uf6XC2lqEQRl8D0BEeSe/zHdMP3qopcKuLoVBzV0PQHGDJr9+fEw/k8qlwskMhkFjRg9At3wlv4VcTD9+8qXC18qGQVBHD0Aa6iS/rBFMPwEJpsJqm4dBUzcPQJyNI792qUs/+EOmwt1niEErUQ9AiOAiv/vsSj+nYKbCNE+JQVJMD0DzrSG//GtKP42FpsKGHopBBWEPQB+WIL83kUk/taOmwovuikHLeg9Ael0fv8GSSD8TzKbC5MKLQWp/D0D+SR6/SvtHPxDwpsInm4xB+5wPQM5wHb/AHEc/hg6nwoRejUG40g9A0tEcv9b6RT+TOafCFA2OQYIMEED7qhu/B4lEP6Zhp8KGso5BfUgQQJzqGr/0QEM/C4mnwkxyj0HjehBAODIav4MjQj/EpafCvz6QQemuEEAbJBm/R9hAP6VLnMJgRnZBcGkTwKNXLL9J2D6/xZOcwjtdc0FdthLAyv4pv5yYQL/k25zCrJtwQUtWEsDrCym/iatBv+wkncKWAG5BSykSwL/ZKL8rSUK/5mqdwsB6a0FXTRLAk4Iov9OQQb+arp3CYfJoQTCmEsAu2ii//lRAvz/rncLrcmZBvOoSwKocKb8LYT+/HC2ewgkIZEEVTBPA/GQpvwT9Pb8Yap7CbqBhQbSGE8BqRCq/1HY9v9mqnsKUOl9Bp70TwEPkKr+V4jy/DOGewnAbXUEgABTAGZArv1YlPL8gHZ/C9sJaQdwgFMCKRyy/YfM7v1xTn8IYY1hBzkkUwDBKLL89UTu/wY6fwmMVVkFeWxTAfsQsv6xAO7+2x5/CXNVTQTNHFMAXDCy/dkA7v2/+n8KkaVFBCDQUwEifK78vXTu/JjGgwpxOT0FBQRTAY8Aqv/rGOr/hZaDCqSBNQekcFMCOcym/UMU6v4iaoMJlwUpByPcTwPqNKL9/8zq/Ys+gwn2cSEEUzxPABL8nv1M5O78tA6HCHFVGQW2iE8DtGie/VqE7vy4zocLsLkRBA2QTwCnmJr/rgDy/2l+hwnT9QUESOBPA8vUmvxk2Pb8hkqHCQr8/QTYHE8BSTSe/PR8+vyDEocKmuT1BTv8SwI2uJ79baj6/TPChwr6mO0HE1RLAtx4ov0tCP7+QI6LCbpA5Qbm3EsDJzyi/MApAv01HosK8hTdBx4ESwCgeKb+MBUG/XHqiwrZtNUGkSRLAD0Epv4f2Qb+tpKLCxXEzQWMiEsCQhym/abRCvxrSosIuZDFBqdYRwFgtKb+nu0O/m/SiwtxzL0H9ixHAT4oovxGdRL8lH6PC02UtQf5REcAQMSi/WV1Fv6BRo8JUeCtBwgcRwKrnJ7+VZka/DnujwtmeKUGnvRDAEFMnvx5MR7/lnqPCQa0nQUeSEMBLuSa/1bJHvwvNo8JEpyVBmE4QwATrJb+NYki/iv+jwhXPI0G7LRDAs1slvyCjSL++KaTCMOIhQcj4D8CepCS/lCFJvyxRpMIg/x9BiMwPwA8lJL8nl0m/ZHakwlYOHkETvA/A/F0jv/x5Sb9hoKTC7jEcQTKID8AwwiK/MQBKvy7DpMIMURpBNlkPwPSXIb9qLUq/vu+kwvyJGEGzRQ/AOKQgv9MFSr80H6XCG74WQU1BD8BwLiC/Yd5Jv6JApcLu3hRBgkEPwO0qH7+sX0m/2milwqIcE0EtGA/Afiwev5GJSb9SkKXCmj4RQcciD8DtcR2/cARJv1K0pcK4aA9B3A0PwHCoHL/h9Ui/2telwjGhDUGdzw7Az4wbv5RjSb8Z/qXCfMkLQWK4DsBqZRq/YS9Jv/gepsLw2AlBZ3YOwPocGb8ylEm/gUCmwuL7B0G4MQ7Ak9QXv6MCSr9CYKbCcE0GQeHfDcBJrRa/l7RKvzmIpsLykQRBnY4NwO2VFb84a0u/jq2mwgDCAkHrOg3AAnIUvzYkTL8XxabCptwAQU25DMA7AxO/MWxNv1rtpsIBm/5AckkMwPLsEb+YmU6/sgynwtAU+0CpywvAv5cQv+HcT79/LqfCQlP3QPY4C8DsdA+/ZoxRv1VEp8LQgfNAq5wKwEsgDr8vR1O/Jm2nwo2170BxDQrAmVMNv60VVb+0k6fCzR3sQOpqCcA4LQy/QQFXv9awp8JXr+hA79AIwE4kC79p2li/ENCnwia05ED7LQjA+kMKvzLuWr948qfCoD3hQLWIB8AAPgm/6/Zcv7YEqMIeqd1APd0GwEQNCL+8AF+/Vh+owgrr2UCyPQbAfe4Gv1XkYL9qQqjCHd7VQN+QBcDDnQW/6eBiv4hVqMIWCdJAJwEFwLyfBL/NlmS/z3eowv9pzkCXcQTAy4oDv3U+Zr9SlqjCjmrKQDnwA8CLpwK/aMlnvzKuqML2eMZAjG4DwLRYAb+gFGm/ssWowvG1wkAb+gLAczIAvwxBar9Y6KjCm0u/QFj4CMCz0P6+ZPxqv+T7qMLK+rpAvc4IwJFH/b6PrWu/KBOpwgcpt0AYlwjADkn7vpWKbL8mL6nCAgGzQKhsCMA4TPq+L7Zsv2tLqcKzaa9ApkUIwOkV+b4UJG2/bFqpwhsqq0CUNgjAbID4vvlobb+1S5/CVfFcQZB+E8DrfCy/KpQ+v6CTn8JwCFpBu9ESwAkuKr9DQEC/ntufwnJIV0F5cRLAAzkpvxdTQb+FJKDCoK9UQV4/EsCk+ii/nP9Bv3ZqoMKIK1JBXV4SwJ+UKL/cVEG/+K2gwpakT0G/sRLAYtUov5okQL+C6qDCnCZNQeLwEsD6BSm/RT4/v0UsocLlvEpBi0wTwFE7Kb+S6D2/BmmhwrdWSEHnghPAFBAqv7tuPb+sqaHCGPNFQQC2E8BrqCq/nOY8v7jfocKv0UNB0fgTwN5UK79mKDy/uBuiwl19QUFcGxTA4g0svwXwO7/dUaLCdSE/QYRFFMB2GCy/lUw7vwGNosJh2DxB51kUwOiULL+7MTu/08WiwmqdOkHvSRTAweArv58iO784/KLCijc4QSA4FMAUeyu/BT07v6Iuo8KwIzZBm0cUwCGqKr8BpDq/DGOjwiX9M0EXJhTA9GApv8eYOr85l6PCj6UxQbsCFMAOhyi/GMU6v63Lo8LWiS9BZ9wTwDW+J79BBDu/4/6jwsxLLUHPsBPAbSUnvxxtO790LqTCoS8rQWNzE8Ae+ya/XE08v55apMIwCClBqkkTwJITJ7+T/Ty/cIykwlTTJkG8GRPAmXYnvyXoPb/AvaTCkNkkQX4SE8Dr5Se/4TY+v1TppMJ20iJBIOoSwKpkKL+iED+/rRulwlLHIEFFzRLAsCYpv03bP7/jPqXCwckeQZGZEsAVhym/1NVAv0dxpcJsvRxB3WMSwB29Kb/JxUG/BpulwjTQGkHcPxLAWBkqv4OAQr+xx6XCoc4YQb32EcAuzym/5YRDv2HppcJc7BZBALARwJFCKb8uYUS/WhOmwiTtFEE5ehHAqf8ovxMbRb/SRKbClA4TQTg1EcC0yCi/HRhGv2ltpsKGQxFBDe8QwCNIKL+b90a/bZCmwq1iD0FayhDAtcEnv9xMR7/hvabCe2oNQaqLEMBFBCe//vBHv8jvpsLAogtBP3AQwP6HJr8GJUi/OhmnwtnFCUGKQBDApOElv+CWSL/IP6fCEvMHQdoaEMAEdCW/+/pIv2Jkp8LLEwZBpg8QwAm8JL9y0Ei/lo2nwrxEBEGV4A/AJC0kvxRKSb/ur6fCMnUCQXq2D8DDEiO/ZWxJv27bp8KWvgBBH6gPwOYpIr82Nkm/KgqowjIG/kDepw/A0MMhvxwGSb8MK6jCLWj6QLCrD8C+yiC/8X5Iv4dSqMI8A/dAsoYPwI7ZH7/+nki/hHmowlhp80AalQ/Ary8fv3QTSL/+nKjCqtnvQO6DD8Cadh6/yv5Hv/q/qMJeaOxAaEgPwNBqHb/Maki/0OWowqjd6EDNMw/A5lYcvw03SL/2BanCsxrlQLb1DsBCIRu/YZdIvz8nqcJkgOFAtLMOwOLrGb9kBkm/okapwoBD3kAkZA7AaNgYv+66Sb/ubanCVezaQBcVDsDG2Be/YHZKvxOTqcIWbtdAEMQNwEjLFr/2MUu/MqqpwgTE00BiQw3AyW4VvyuCTL9b0qnCr8HQQMTWDMD9dBS/arNNv5DxqcJ0Xs1AmVoMwCE0E799/U6/HxOqwjTByUB7ygvAMCcSv9iwUL8PKarCeA7GQNQwC8AO7RC/NnJSv5xRqsKEY8JAYaQKwII5EL9fRVS/KniqwtDpvkB1AwrAjSUPv6s3Vr/klKrCs5u7QN5sCcB0NQ6/uBNYv6m0qsLwwbdAtMsIwEJqDb/sLlq/MteqwtxstEAOKwjALXkMv+YzXL976arCcviwQGCDB8C1XQu/GT5ev2QEq8JoXa1AY+cGwNxRCr+eIWC/tCerwrJ0qUAyPwbAuhEJvwUZYr8GO6vCUcGlQO2yBcCeJAi/9M1jv3Bdq8KiQ6JAEygFwF8gB7+kb2W/03yrwpxrnkCWqwTA2UsGv1byZr+olKvCNJuaQC8uBMCWCAW/zTZov7Csq8J0+ZZASL8DwHDyA7+XWWm/us+rwgSzk0DveQPA0TUDvy4Gar8A5KvCloSPQPk3A8DlfgK/4adqv7P7q8Jy3YtASuECwEuEAb+edWu/vxeswuLVh0CvxwLAkAgBv9yTa7+YNKzCMmiEQHmbAsA4dwC/lvFrvztErMKDTYBANYMCwP4vAL8tKmy/cHmkws/AGkEnsDbAIwJzv8VAkL7zk6XCXBcYQaVUNsA20nO/zFaTvtetpsK+jBVBNec1wAFPdL8e5pa+CsmnwqQHE0GqeDXApsB0v3N9mr775qjCcIwQQZ1WNcDyyHS/1JCbvqn9qcKMAQ5B3zQ1wNwHdb/sspy+MRSrwjqRC0GkRzXA6+J0v+QQnL67KKzCB0MJQXuCNcAer3S/Zimavuw9rcKj8gZBLpU1wG9edL/tepm+1U6uwpbRBEG3yjXAVSt0vyS/l77yVq/CDegCQUZfNsCdSXS/HyWTvkdhsMLf/gBBtc02wH4Idb9w6o++HGSxwvvS/UCNaTfAvEB1v7Qdi77qabLC/yP6QFb9N8AlYXa/EtCGviNus8JaV/ZAmbE4wAv6dr/1V4G+bWW0wq7J8kAKfTnAAoN3v+RCdr61WbXCWnjvQKhSOsDXFHi/jTVpvs9MtsJQNexAKxM7wGjCd79ZEl2+Bj23wsTm6EC5wDvAkrp3v45DUr58KLjCbgnmQA5jPMAgAHi/WUlIvkoVucLyLONApPo8wCRIeL+q+j6+se+5wsi+4ECBZj3AljJ4v0NAOL4i17rCXh3eQAnmPcA6QXi/6VswvjKuu8JcSttAkkA+wMVpeL9wzCq+64m8wo9V2UDrmT7AkbV4vx9bJb4OV73C90TXQAbsPsBhAHm/gFsgvrcwvsLEUNVAyTQ/wMVOeb9Y8Bu+rva+wlpz00DLdD/A4LF5v5YVGL5OwL/CSdDRQP+cP8BiB3q/+68VvuR+wMJKndBAFApAwAC8er8DHQ++PUfBwkYBz0B9P0DArv16v0jdC766/MHCaqDNQD+IQMCKdXu/P3cHvoS8wsJlh8xASdhAwCtyfL98vQK+5HfDwqxoy0DXKEHAIUx9v67c+72RLMTCFUfKQD6HQcCV1n2/H1bwvdjOxMKvkMlAgv9BwA3mfr87yeG9P4LFwhagyEByWULAl49/v0rQ1r28J8bCseXHQN38QsDnDIC/8o7CvS/MxsLXVsdA0nVDwEtdgL/Nn7O9JWTHwtcQx0BlE0TAmZ6Av1AGoL24BcjCz6nGQP7BRMBBzYC/XzeKvUuXyMI6+cVAv3xFwErzgL+br2W9WiDJwoq/xUBLHEbA/NiAv5WMPb0Kt8nCEQTGQOboRsBR0oC/jjEKvRQ/ysK0GcZAQtxHwLzXgL9ET5q8t8PKwhpRxkBnnUjAptaAv9in5bsERcvCTK3GQCedSEBcuYC/qQzmOxbGy8If9sZABpNHQNJ/gL9by748nE3MwstKx0D2vUZAlqaAv/3dFD3DyszCOz3IQBDARUB9coC/lV1UPYA+zcIecMlAOMNEQCY4gL/Nv4k9CaXNwmgQykDIx0NAcux/v4AUqT2YF87CufzKQNLAQkCpXH+/dsrJPeOOzsJqy8xAlN5BQI65fr8P1OU9lvHOwkQUzkB8+kBAgIF+vzohAT6CY8/CZIXPQOnsP0C3oX2/Z8ERPqm9z8KC2tBAEB8/QHvKfL8BXx4+thXQwqnF0kDhLT5AIh18vyw4LT5agdDCAr/UQLN4PUAVU3u/8EE4Pm3k0MLnm9ZAFaU8QPSXer+6MUU+yTjRwv/f2EB44DtA8Pp5v8Q4UT6gj9HCeZvaQIYcO0D4QHm/wCRdPqD60cKIE91AqpI6QNnneL/2l2U+70zSwpf/30A+6TlAFEZ4v13gbz48mtLCFB/iQDhgOUCE1Xe/2Td4Pnv90sJE4+RAq+84QBkXd78m334+kj3TwkwI6ECieDhA1Vd2v9/0gj7kk9PChG/qQHYVOECC+XW/ofOFPizT08JbOO1A8bk3QKgQdb8Kj4g+A0DUwvDQ70DdiDdA3qJ0v0D4iT4PhtTC3hzzQLVGN0CfS3S/Uu+LPkzU1MJaSvZATg03QLUsc7+mZ40+EA7Vwnbd+EDS/zZALJpyv1SpjT6BWdXCu9n7QBfKNkB89HG/1ySPPrqp1cIEV/9AbZc2QE5rcb9kj5A+CPLVwg0yAUFYljZAYiZxv/ODkD7TNtbCT4cCQa2zNkBbQXC/KFmPPs+G1sLx7wNBr8U2QCT0b79qtI4+NtPWwn+gBUFp0TZASm5vvz0xjj7uEtfChBwHQRr4NkA1mG+/iguNPkVQp8IsFBJB9QAowM9ZX7/XQf2+rTuowqADDkGHMyfAUn9evyCeAb/IJ6nCcRoKQQqhJsBDSV6/u9sDv9IUqsLVPgZBsjgmwCZ8Xr/TjQW/KQOrwglvAkEkMibAvGRev7WhBb/N66vCDSP9QFdhJsByxV6/ev4Ev7jQrMI4kfVAp6cmwNvTXr9R5wO/lritwoYb7kAYEyfA3O5ev6M+Ar8Jnq7CmpLmQOpCJ8BRZl+/2Z4Bv56Dr8L1Vt9AMYwnwD+sX79CiwC/fFywwq0K2UD4ESjA+0Jgv340/b6VPLHCURnSQKNiKMCGJmG/RyT7vnoQssIy1cpAXtAowFRTYb/7yve+G++ywn0DxEDQHynAg1Biv1XP9b6HyLPCbxm9QI95KcCvLmK/hu7yviuYtMLGJ7ZAXtUpwKAlYr/uCvC+hmK1wialr0BRSirAod5hv1VB7L69MLbC8SSpQG2iKsDitGC/OfLovqT9tsJNUaJAquAqwDYVYL/Ntua+JMa3wnLum0BCGCvAXrVfvyzQ5L50kLjCo22VQJpEK8Dcbl+/kU/jvj9OucLWSo9AcjkrwF4uX784ieO+DQu6wvUciUCiVivAKkBfv9mq4r5KyrrCa4KCQCpFK8Avp1+/1mXjvtOKu8L0lXlAhFgrwENMYL+6GuO+EUC8wk7lbUA9VCvAP/1gv16Q474q/7zC6O5hQHI/K8Bz8GG/xajkvtqpvcKmoVZAzRorwBS8Yr9vLea+WWC+wvYjS0A25yrAgGljvxwc6L6jCr/CWkNAQDbiKsCGemS/NcfovrbAv8Ko7jRAYLAqwObkZL/Xieq+cmLAwkKmKUClfyrAtjVlvyI47L5IDsHCAbEeQLZfKsAXD2a/66Ltvmu9wcKiARRAEDQqwMPHZr8mXe++omjCwhZVCUDNEirA/0hnv0+p8L7Q/MLCTRj+PwANKsDX3Ge/PiHxvr+qw8IIKOg/VOwpwHM/aL+KWvK+RlDEwiVm0z9K/inAPJFov5Lx8b7S8sTC+te+P6XzKcDKG2m/iozyvoqHxcLH46o/APspwLF9ab9egfK+ziTGwnkalj/BFirA8d5pvyLQ8b4uu8bCmlWBP00rKsD/Omq/nlbxvgRJx8KUFFw/VzwqwJhBar95z/C+POPHwsg5Nz/AayrAW2pqv2ti776eeMjCVF8RP4GtKsB2/mq/eJTtvkf/yMKF9dM+R8kqwIpFa7871ey+sIvJwr+Liz689SrA0lVrvxF0675CHcrC1mv5PeM+K8BCjGu/ZjzpvtWmysJfgwu9KmArwML2a7/GYOi+8DTLwt6fK75PfyvA2d1rv2FY576zs8vCsvSXvtiwK8DApWu/iazlvhMvzMLiJee+hscrwKpRa7+tzeS+3LHMwhHYGL+q3CvA7xhrv0oI5L4yMs3CkqY2v+jOK8B0jmq/Hjfkvq+vzcIz/Fq/bdUrwDyvar+tEeS+5DTOws/Dfb+96CvAdTtqv/c/4742oc7CruyRv2/GK8CHzWm/uCHkviQWz8J6m6G/UcArwAhzab/jKOS+35HPwnzQsr+GiyvAFQFpvzyd5b73BdDCXabEv/ZlK8Acs2i/VKfmvnRu0MIKiNa/QT8rwLpZaL8Gtee+/OXQwic+6b9CEivALDdov1YP6b6QZNHCv1n6v7nOKsCRGmi/WyLrvnDN0cKmAgXAu6gqwAy3Z7/TJOy+PzrSwv7dDsBAbyrAaMVnv7j77b42sdLCRx0XwHkhKsBnSme/fjPwvnwJ08IlCR/AOdspwBXdZr/JNPK+1nHTwiMGKMAhrinADFlmvxxf87511NPCgxAxwFlfKcBGh2W/mHH1vnFB1MLqJDvAvispwCQJZb8T0va+r6nUwoaGQsBjBCnAC4lkv0DO9756ENXCX95LwMnIKMA+v2O/OEf5vjhm1cJ0DVXATpoowKiqYr+nLvq+BszVwrxPXcCKhyjAyrVhv3ZG+r5oMdbC+EVlwK+UKMDW9mC/0Xn5vvyM1sIAP27Ahn8owN1+YL8o5fm+St/WwnLmdsDVYSjA9yNfvx8e+r6GSdfCugGAwNhqKMA2zl6/Xqn5vgSo18Iy7IPAtWAowKLTXb+Gd/m+G/nXwouaiMAlbCjAna9dv5EJ+b5O5F3CJLB7wi5tr78kUEs+/sx6v/hsXcII9XzC3Pmuv1yAWD43lny/GPhcws46fsL3Da+/rAFlPvxKf78YhFzCVIF/wgHdrr80m3A+uaGAv1YYXMI+Z4DCZn+tvwU8fT63gYC/IJ5bwjsQgcJ0tKu/GMyDPndnf79YKVvCpriBwoFGqr8NxIc+YQt+vxq4WsIPY4LCF8Oovy/Cij4nIHy/PERawlIPg8IuFai/D7SLPlAce78E1FnC3rqDwsLMp78hEYs+4Ep6vwJdWcLHaYTC7a2nv9gfiT6HTXm/zPBYwlQVhcLiY6e/5OqGPu3fd79Sc1jCG8eFwk2Ipr8GSoY+d+t1vy/8V8J3d4bCuV+mvyvvhD7OFnW/4olXwlUgh8JdKKW/3kqGPhoxc7+LFFfCsdaHwqTOpL8qcIY+4Y1yv2iiVsLvgYjCbfGjv1QHhz63EnG/vB9WwoovicKRjKO/8p+HPmSFcL80oVXCKdqJwuVHor8wtog+vXBuv8AsVcKsjIrCv4qhv5KiiT62Vm2/ZLVUwsMvi8L8Z6C/nrmLPk7ka78yM1TCmuCLwhO7n79PWY0+IytrvzGvU8Jei4zCcZ2evyGDkD5NJGq/ejBTwro3jcI8jJ2/2CCUPpBcab8Ar1LCmNqNwkI5nL9AsZc+cg5ovyslUsJlgI7CdPiav8wznD7VNGe/86NRwpkjj8LRm5m/XHCgPmsJZr+BEFHCkMOPwsZ3mL/KiKU+ApJlv0SEUMLHY5DCoSaXv/6Wqj7oumS/hgVQwgsAkcJ2t5W/StyvPuC2Y7+Ma0/Cgp6RwiHzk7+torY+V4Ziv5LZTsLoM5LCmB2Sv4w8vT4IHWG/CFdOwmXNksKCIJC/rKfDPsBPX7+Mr03CB16TwmeCjr/ySco+BkNev2Q3TcId8pPCIaSMv0zS0D59qVy/35NMwtOIlMJft4q/YP3WPlbPWr8tEUzC4hCVwlpgiL/a+t0+al5YvzqCS8IqoJXCccOGv8aN4z5G4la/IPFKwioulsL5F4W/ZaTpPoRpVb8IWUrC3bWWwtb+gr8H1u4+INVSvzfHScJdPZfCV/uAvzXV9D5CnlC/DxdJwiDGl8I2yH6/0Z/6PrMiT7/kmUjClkiYwgm4e7947f8+TJxNv2r4R8IlzZjCF4R4vx6IAj904ku/1mZHwsNHmcKNdnW/vUgEP9zbSb/OukbC38mZwglHcr/uagY/a+RHv10tRsJGRZrCdRJwv6wtCD+Uqka/HKBFwqG+msLdCG2/cwYKP1asRL8W9ETCbD+bwkA/a79wgAs/NrBDv85XRMKNtJvC0Ytpv5iLDD9HkUK/Ss9DwhQonMKhE2e/ASkOP5P7QL9aKkPCGqecwh27Zb/+GA8/ryRAvz6kQsKZFp3CLddkv5+XDz+0hj+/nhBCwuKKncJuAmS/FGcQP7MdP78kgEHCP/SdwnMjZL+GFRA/dRY/v8PVQMLxZZ7C3Q9jv4taED8JLz6/YkdAwo3VnsIMQWO/FC4RP8TCPr94tT/CikSfwquFY7/PARE/oe8+v2AhP8KXs5/CUkBkv3nfED+tkj+/gYQ+wh4joMJ/oWS/lK4QP8LYP79A8D3C242gwsImZb+G3RA/jW9Av/VrPcKp+6DCFvpmv6bCDz9/qUG/re48wvNkocJT7me/7pMPP2x+Qr+YcjzCmtChwjWVab+Qrw4/N6dDv8HhO8KbQ6LC3klrv7FGDj/qGUW/Llo7wsKsosJw42y/uGsNP/A5Rr+v0jrCsBqjwmsTb7+K6gw/xxhIv55TOsJUjqPCogpxvw3yCz+ChEm/cMo5wv35o8LJo3K/+CMLP9GpSr++ODnCUmqkwv6TdL9tEgo/2ABMvz7TOMLx3aTCrEN3vxWECD/w0E2/8Ew4wgdSpcIqSHm/5w8HP8kFT79jwjfCmM6lwoZle7+LGgY/IJVQv8IwN8KpPKbC7C19v3oBBT+TvVG/1K42wuawpsJTrH6/aLQDP1SAUr/eZTbC8iinwjRSgL83RAI/26VTv3bGNcJSpKfCcTOBv/HJAD+dj1S/HmI1wnAcqMJ5EYK/8a7+PrV1Vb/m3jTCPo2owhHsgr/Dtvs+BU5Wv9djNMJPDKnC7IODv2So+T7B41a/1f2kwpdJfEF4icC/qDV6PZsffr/z7qTCpmF7Qc2XwL8+so49fyWAvzTgpMKmiHpBNt/Av9C1mD2BA4G/oNakwrDXeUEzJMG/aaedPZqTgb8SzaTC7jV5QT/swL/CIqM9M62Bv/nEpMKconhBJqvAvx1rpj1dnIG/nrqkwpYJeEGnfsC/M0OpPcuZgb+/taTCWIt3QcllwL8Qmaw9cbKBvxKtpMIgGndBW2LAv1zZrD22soG/A6mkwgmgdkFmuMC/8qWqPTfpgb+woKTCQgF2QfzbwL9Qj6g9Fu6Bv8icpML5onVB48jAvy62pT30r4G/TJekwkEjdUFbacC/C0CmPSVXgb8yk6TCjrx0QWpxwL/59KM9Aj2Bv8uQpMJWaHRBhAXAvy7npj3o+4C/EIqkwqHmc0GBD8C/e4SlPVbxgL/Ph6TCsZZzQXrNv7+8XaY9WbuAvzSBpMKRR3NB4su/vwL8pj3xwoC/Vn6kwrXeckEcn7+/WpinPeSegL+GfKTCc4pyQaeKv79nJ6o9SbCAv9Z2pMJINXJBmDW/vxSfrT0AjoC/T3Skwtz1cUGCSr+/9bevPSnCgL/icaTCxIpxQW7vvr/Y+7U9AcOAvz5rpMKgQXFBK9C+v3guuz078IC/FGikwlUYcUGXLL6/Ib3CPVm6gL/bZqTCc81wQVjhvb/yEss9rOiAv7NipMI2mHBBrlS9v+q00z3C2IC/AFqkwrVOcEFHr7y/xEfePYHLgL8oWqTCPAtwQZPtu7871Oo9YL2Av91apML02G9B4Om6v38k9z0WaIC/O1KkwsGPb0FxyLm/IwIDPkAYgL92T6TCDIJvQUWHuL/hywo+v2F/v3tIpMIDSG9Bxgm3v8xAEj51BH6/ckmkwsIcb0EtvrW/qUEZPuzvfL+YRaTCQOhuQR+FtL8HrSA+5xR8v+ZCpMLHrW5BG+Wyv7y3Jz4SVnq/EEKkwjh0bkEzZrG//7suPmnWeL+bQ6TC6VhuQWHRr78/gjU+Wx13vyJDpMJdKG5BhGGuv1t9PD6YtnW/6UCkwgLobUGC6Ky/Xf9CPvUjdL90QKTCg6dtQXNFq7/4nEk+30Nyv288pMLYZ21BCtSpv7SNUD6o03C/dj2kwscsbUE+fKi/rlZYPv68b7/aQqTCnwRtQXzfpr959V8+OBZuv+5GpMLU1mxB2Iulvx7BZT7XoWy/50WkwmGPbEHJF6S/HpxrPk7xar/KSqTCnmFsQV4No7+M/HE+niJqv8xKpMJQJ2xBQrahvxgHdz7igGi/MUmkwvXsa0Fu+aC/WVJ6PhizZ78ISqTCLK5rQXAboL/wE34+VrtmvzpRpMKkaWtBJVKfv4E7gT5gB2a/FkykwiEfa0H0156/7uGCPtG2Zb/ITKTCJc1qQdBqnr+QJ4Q+/FtlvxZKpMIiiGpByiuev8gQhT5vN2W/+k2kwhJqakEnPZ6/lyGFPlheZb+lTaTCXjFqQS5Mnr+Xo4U+Gallv9xHpMJMAmpBZYGev3xohj5hVGa/t06kwuzVaUGC956/U8iFPj38Zr9aUKTCCbJpQYFsn7+Kh4U+WMRnv0hTpMI1d2lBvvCfvwIJhT7Jk2i/BkukwntKaUGXTqC/WnOFPiVuab9QUqTCyARpQcyQob/TAIM+HvZqv85SpMLMymhBpk2iv33Jgj5iTmy/QEukwlS/aEF8P6O/EmKBPkCebb+XUKTCFnpoQXhnpL85DYA+Dl9vv9hRpMIGS2hB0nmlv7AcfT4j53C/Tk6kwgINaEEfW6a/Tt17PpJkcr9YU6TCM+lnQdqKp7/C13c+LvZzv5ZTpMKUjWdBAW2ov189dT7gNHW/8EqkwklrZ0GKbqm/vqFyPjyydr8ISqTCnxNnQXdzqr/9v28+nyl4v7BLpMI06mZBOkarvzFPbT7FU3m/Ykikwr+fZkF5Fqy/jhJqPtNQer8YRaTC8DtmQV7ZrL9CDmc+Qz57v8ZDpMKlFGZBu1qtvyxVZD7Htnu/ZkikwiWqZUHRCa6/Xh5gPgs+fL8VQqTCdU9lQTCyrr9mSlw+/8p8v/ZBpMJ91GRB7/+uv1SFWj69C32/b0Wkwu6gZEGrrq+/LwpVPuhNfb/yQKTCPB5kQZvvr79sX1I+AUV9v+aYpcJr255BbJI/vyEmMj84eSq/3IOlwqvHnkFQ5EK/gmo0P2CjLr/IcaXCra+eQfrLRb/orTU//Qcyv9VlpcKzoJ5Bx6VHv3EXNj8ZDDS/8lqlwrCankGcdUi/3Bc2P0bcNL+MUKXC3ZOeQRaqSL+5vDU/K+s0vwBGpcL9ip5BsQdIv0JvNT/KKDS/4j6lwumQnkGnYEe/gj41P/RtM7/UOaXCWpqeQTMxRr/KBDU/7Ccyvy80pcIen55BRnNFv23fND/QWzG/Ai6lwiWcnkHmzkO/xfk0PyXFL7/kKqXC/6qeQXElQr+W6TQ/YRkuv+4lpcIlrp5BmDZAv66ENT/obSy/AySlwuWynkFwKT+/z6o1P1VzK7+KI6XCfr2eQYzRPb/IPDY/WVkqv7ofpcKesJ5BzMk8v4FvNj9JaSm/Ix+lwhy6nkEImju/mTo3P4iMKL9HGaXCg8WeQS0XO7/iBzg/VVoov7UZpcL8wZ5BiBY6v2JHOT8w1ye/vxmlwmHBnkFUWTm/WHo6P+CQJ7/iFqXCBcSeQchTOL8FCTw/sSQnvx4WpcIDz55BEfw3v8JcPT99TSe/TBalwhzEnkEuCDe/6lY/P6QYJ7+sEqXCpcOeQYeONr/xFEE/s0Unv7ITpcJm0p5BzyU1v8PjQj+5iCa/rhSlwnzNnkFKbTS/u5lEPzdxJr9FFKXCANmeQXo9M78TTEY/Ld8lv08QpcLDzp5B5lcyv/isRz+ceCW/shOlwpDRnkGIdTG/DBVJP8YWJb/aF6XC4NSeQbDqL7/iR0o/gvYjvx0VpcKOzJ5B5nMuv7KOSz9d8CK/kBalwonYnkGxCi2/8XNMP1LUIb/cFKXC0d6eQWkRK79MS00/OiEgv5YYpcL62Z5BfoIpvyH4TT9Wyh6/zRilwmzVnkE0LCi/y5FOP6+lHb82F6XCVNOeQWn/Jb+eBE8/npsbv34bpcIV0p5BoQ0kv2KcTz+d2Rm/7h2lwrbcnkEA6CG/XlFQPxvtF7+qHqXC1tyeQRoQIL9i7VA/UkYWv0khpcIZ4p5BMMYdv7V8UT8nKRS/NyKlwrLmnkHXfBu/hiBSPzsTEr8yIKXCrdmeQQpdGb/6+lI/szcQv6ogpcIj3Z5BhIkXv3DQUz84pg6/8yalwhLfnkGUbhW/wcVUP2HWDL/eKaXCn+CeQdNkE7+DZ1U/E/4Kv0gqpcIG3p5BdcsRv0q+VT+ifwm/KiylwlHcnkF5iRC/QWtWP09xCL92LqXCreSeQTV/Dr9JmVY/13YGvzQqpcKe2p5BAskNv1D2Vj833AW/cSulwt3enkGFWAy/N7lWP/5cBL92MaXCBd2eQRE/C78H4VY/ClEDv/4rpcLn2J5BquYKv+YNVz/pBQO/FSylwlzJnkG0hwq/i+FWP4CbAr9VKaXCMMGeQaw+Cr+EzVY/o00Cv2wppcLZvZ5BMxkKv+xWVj/CBwK/Ji6lws/BnkE0pQq/FwJWP7B6Ar8dJaXC87yeQdnjCr9QuVU/ZaQCvwAlpcKdtp5BkYMLv655VT+IMAO/NCqlwvHDnkEX8Au/RlhVP3mSA78HK6XCtbaeQVzODL+qH1U/Xl4Evw0lpcK7up5B2VINv5t6VT85+wS/viSlwo+pnkEjdQ6/duVUP1fwBb8EJKXCApueQT5xD7+RTVU/JggHv2ocpcJ3oJ5Bta4Qv0omVT8MOAi/hyKlwoSlnkH69RG/iUlVP8aHCb+SIqXC8JmeQSQvE7/tcVU/oMsKvzgfpcJ+q55BazsUvwkiVj+7Cwy/SCClwoGUnkGFqBW/6OpVP3FoDb9yHKXCUYOeQeitFr8tH1Y/EX4OvzgXpcIxhp5B+/QXv0gWVj+Cww+/NhSlwih3nkHwPxm/7h9WPwcTEb+GFKXC5XeeQZ6FGr9wFFY/OFcSv8MQpcIrgJ5Bt4Qbv7LKVT8sQRO/0hGlwsNrnkGN0Ry/QTxVP8xjFL+yE6XCQlmeQdoAHr/PvFQ/KW0Vv+UWpcJAX55B9FUfv3elUz+caxa/rw+lwng6nkGQqSC/5PhSPy2KF7+4EqXCqxWeQaXCIb/GG1I/dF0Yv6oWpcJ3HZ5B6gIjv4DbUD/eNhm/mxSlwijynUEVJyS/dPpPP3ASGr+e1nvCmG1mwjPWxr6ATWg/vQ3AvkaLesJG7WbC+XjMvnYTaD8FjcW+gD15wrdqZ8IwRtO+h6BnPxgfzL4k9HfCFOdnwgDG2L788GY/w0zRvt6vdsJbZmjCV8DcvrylZj/kIdW+VGJ1wojraMLbB9++9XFmP3lQ176oF3TCKm1pwrfs4L6PgWY/3DrZvgbVcsIW7mnCjfvhvi2yZj/qXtq+/JJxwnR1asIOPeS+jfxmP4HB3L4YTnDCjvZqwniA5r7hzWY/2O/evloIb8Jih2vCornnvsqhZj/+FOC+otFtwnkCbMLEdOi+MGtmPxa34L7UkGzCzoNswqRI577+omY/eKTfvupUa8I4Cm3CKCjpviIrZj/zTOG+0Bxqwt+BbcKyxee+tr5mP2Iu4L6E62jCaQtuwgms576GwGY/iBXgvj6uZ8L9hW7CnG7mvvBxZj/hs96+Tmlmwj/+bsJ4KeW+xEtmP39d3b40NGXCVHJvwhZ74r7eK2Y/3aHavt4CZMIg82/CUajgvrnnZT8Hsti+sNFiwmxYcMKC/t2+dellP/QL1r6yomHC5tFwwgyV274FuWU/v5DTvlhnYMJKQ3HCyOXXvhBcZj/3LdC+uDFfwlm1ccKs/9O+KxZnP2WdzL4ICV7CqhpywkwP0L6TTmc/z8vIvn7OXMK+g3LCVpLKvhAwaD/FtcO+YLVbwhDjcsLR5MW+ycVoP79Nv75ud1rCP0ZzwhTCwb6puWk/kJK7vnhTWcJ0onPCwE2+vrwWaj8/Sbi+0jdYwv78c8LSFLm+mrdqP1hYs75rEVfC2Fd0wmFctL6e8Gs/4RuvviL1VcLKqnTCjNivvpuYbD9R3aq+6OFUwtQBdcJ57Kq+kTxtP9A0pr4Qv1PCO011wlfepr5RBW4/B3OivkK4UsKOo3XCz6KjvvZRbj8ZWJ++bJ1Rwpn8dcKffJ++tPpuP59xm75QnVDCmDx2wvODm76ZxW8/OsGXvi6aT8LUj3bCTCaYvnolcD95iJS+YJZOwiXkdsKML5a+nbZwP2bBkr6+lU3CvSh3wjK8kb4aInE/dneOvjeRTMIIanfC8MCOvn2ecT9fpou+FZJLwiS6d8LUL42+hlhyP5FMir7rqUrCeP13wtKSi77TvXI/ts6IviylScJuRXjCIbiJvraYcz/3Moe+QrFIwmeGeMJhO4i+ushzP5DFhb6PwUfC3Mh4wsA8h747RXQ/q+mEvnjMRsLuCXnCE4yGvubfdD/jYoS+lO9Fwvo/ecI/6YW+qjh1Py7Yg77E80TCwIZ5wnDxhr5VzXU/uwaFviEgRMJ8wHnCs4eHvtq4dT8ll4W+8UBDwjf6ecKZyoe++1t2P2IFhr7LXELC/Ed6wkZ0ib4+2XY/ZtCHvjqLQcIWg3rCha6Lvibmdj9hDoq+Er1AwoK8esLyk42+mOV2PyH0i76c9D/Cqut6wnWukL7rP3Y/1eCOvkklP8I/LnvCdfORvkv8dT++EpC+blE+whZje8JmepW+Ulp2P1q2k76biT3Ct6x7wtvqmL4x1nU/kgGXvuXFPMJO5HvCjJ2cviEGdT/Zdpq+4A48wk4pfML1bqC+4mx0Py8anr5XXDvCjld8wnulo77UCXQ/pzKhvvqjOsKio3zCBuunvjmBcj82+aS+AvM5wsTbfMIKvqu+xJtxP29/qL6FTTnCRBp9woImr75+XHA/4nmrvq2aOMKhZH3Czr2zvpUnbz+foq++b/83wp6hfcLzSre+u55tP+mfsr4xTTfCF+p9wpouvL5LRmw/L/+2vmeuNsJzPX7CrxrAvjtIaj//JLq+Ugk2whOEfsJDfsO+3aloPxfivL4aXjXChr9+woCFxr7GgWY/SAm/vinNNMJVDn/CTzbLvtEuZD8SvMK+6C40whhdf8JBls2+wPFhP0QpxL5KezPCUZt/wkcn0L5MEWA/mufFvqPVMsJC7n/CinvSvt04Xj+0ace+DToywuYlgMI8J9W+YGpcPwxByb6mwjHCSEeAwlOZ1r6odlo/UNTJvob/MMLCdoDCcZLYvhrMWD+bBsu+FG8wwlaigMKGddm+wEhXP/k6y75+sy/CQ8aAwlkW2r6YzlU/WzLLvhgvL8K584DCBTTbvpWBVD8EtMu+ntZ7wphtZsIz1sa+gE1oP70NwL5Gi3rCRu1mwvl4zL52E2g/BY3FvoA9ecK3amfCMEbTvoegZz8YH8y+JPR3whTnZ8IAxti+/PBmP8NM0b7er3bCW2ZowlfA3L68pWY/5CHVvlRidcKI62jC2wffvvVxZj95UNe+qBd0wiptacK37OC+j4FmP9w62b4G1XLCFu5pwo374b4tsmY/6l7avvySccJ0dWrCDj3kvo38Zj+Bwdy+GE5wwo72asJ4gOa+4c1mP9jv3r5aCG/CYodrwqK5577KoWY//hTgvqLRbcJ5AmzCxHTovjBrZj8Wt+C+1JBsws6DbMKkSOe+/qJmP3ik377qVGvCOAptwigo6b4iK2Y/80zhvtAcasLfgW3CssXnvra+Zj9iLuC+hOtowmkLbsIJrOe+hsBmP4gV4L4+rmfC/YVuwpxu5r7wcWY/4bPevk5pZsI//m7CeCnlvsRLZj9/Xd2+NDRlwlRyb8IWe+K+3itmP92h2r7eAmTCIPNvwlGo4L6552U/B7LYvrDRYsJsWHDCgv7dvnXpZT/0C9a+sqJhwubRcMIMldu+BbllP7+Q075YZ2DCSkNxwsjl174QXGY/9y3QvrgxX8JZtXHCrP/TvisWZz9lncy+CAlewqoacsJMD9C+k05nP8/LyL5+zlzCvoNywlaSyr4QMGg/xbXDvmC1W8IQ43LC0eTFvsnFaD+/Tb++bndawj9Gc8IUwsG+qblpP5CSu754U1nCdKJzwsBNvr68Fmo/P0m4vtI3WML+/HPC0hS5vpq3aj9YWLO+axFXwthXdMJhXLS+nvBrP+Ebr74i9VXCyqp0wozYr76bmGw/Ud2qvujhVMLUAXXCeeyqvpE8bT/QNKa+EL9TwjtNdcJX3qa+UQVuPwdzor5CuFLCjqN1ws+io772UW4/GVifvmydUcKZ/HXCn3yfvrT6bj+fcZu+UJ1Qwpg8dsLzg5u+mcVvPzrBl74umk/C1I92wkwmmL56JXA/eYiUvmCWTsIl5HbCjC+Wvp22cD9mwZK+vpVNwr0od8IyvJG+GiJxP3Z3jr43kUzCCGp3wvDAjr59nnE/X6aLvhWSS8IkunfC1C+NvoZYcj+RTIq+66lKwnj9d8LSkou+071yP7bOiL4spUnCbkV4wiG4ib62mHM/9zKHvkKxSMJnhnjCYTuIvrrIcz+QxYW+j8FHwtzIeMLAPIe+O0V0P6vphL54zEbC7gl5whOMhr7m33Q/42KEvpTvRcL6P3nCP+mFvqo4dT8u2IO+xPNEwsCGecJw8Ya+Vc11P7sGhb4hIETCfMB5wrOHh77auHU/JZeFvvFAQ8I3+nnCmcqHvvtbdj9iBYa+y1xCwvxHesJGdIm+Ptl2P2bQh746i0HCFoN6woWui74m5nY/YQ6KvhK9QMKCvHrC8pONvpjldj8h9Iu+nPQ/wqrresJ1rpC+6z92P9Xgjr5JJT/CPy57wnXzkb5L/HU/vhKQvm5RPsIWY3vCZnqVvlJadj9atpO+m4k9wrese8Lb6pi+MdZ1P5IBl77lxTzCTuR7woydnL4hBnU/2XaavuAOPMJOKXzC9W6gvuJsdD8vGp6+V1w7wo5XfMJ7paO+1Al0P6cyob76ozrCoqN8wgbrp745gXI/NvmkvgLzOcLE23zCCr6rvsSbcT9vf6i+hU05wkQafcKCJq++flxwP+J5q76tmjjCoWR9ws69s76VJ28/n6Kvvm//N8KeoX3C80q3vruebT/pn7K+MU03whfqfcKaLry+S0ZsPy//tr5nrjbCcz1+wq8awL47SGo//yS6vlIJNsIThH7CQ37Dvt2paD8X4ry+Gl41woa/fsKAhca+xoFmP0gJv74pzTTCVQ5/wk82y77RLmQ/ErzCvuguNMIYXX/CQZbNvsDxYT9EKcS+SnszwlGbf8JHJ9C+TBFgP5rnxb6j1TLCQu5/wop70r7dOF4/tGnHvg06MsLmJYDCPCfVvmBqXD8MQcm+psIxwkhHgMJTmda+qHZaP1DUyb6G/zDCwnaAwnGS2L4azFg/mwbLvhRvMMJWooDChnXZvsBIVz/5Osu+frMvwkPGgMJZFtq+mM5VP1syy74YLy/CufOAwgU0276VgVQ/BLTLvswfYMLbXIHCnpE/vxd/Oj+Evy2/PuxewkDOgcJSUUK/BI46P+eEML8auV3C5D2CwhycRb9zZTo/+sIzv0SFXMI6rILCZ7hHv+j3OT/GtjW/sFpbwvobg8JxUEi/nO45P3FMNr9GIFrCfI6DwomhR78h9zk/Yp81v8zqWMIf/oPCFV9GvzdyOj+WjDS/47dXwihthMLyU0S/ZUM7P3vSMr+kg1bC8N6Ewg7nQr+vKjw/IsAxv41LVcIqTIXCEmJBv/65PD/2cTC/SBFUwm7DhcKuqz+/jSk9P6XlLr/I4FLChi2GwpbLPb81hj0/figtv1ilUcIjm4bCV+w6vwpQPj/Nliq/TWpQwiYKh8I45Tm/nmg+P3iZKb9qMU/C+m+Hwg1FN7+xkj8/nmsnv5r9TcIo34fC1VE1v5QRQD/GqSW//L5MwndGiMLOIzO/SkZAP8WSI78JckvCU6uIwhrNML8csEA/12chv9IwSsIqDonC9c0tv6A7QT9rox6/3vNIwil2icIMRyu/jMJBPwlVHL+ztkfCqc2JwsRHKL+prUI/TbMZv6l2RsKWMIrCokglv2qbQz+JEhe/LCZFwnqOisIHsCG/TIJFP9UsFL9I2EPCle2KwjXyHb/Cmkc/CS8Rv1iTQsIcQovCmu0ZvwtwST/G0g2/gDtBwjqai8IMjhW/wdpLP0tHCr+s/j/CYeqLwrJNEb8mFk4/sMYGv1SePsLjOozCJ4cNv8ORUD9+ygO/01A9wl+JjMJ7+wm/H4JSP3XbAL+QEDzCrtSMwnCxBb8sglQ/4GP6vha+OsIoH43C8JoBvyIZVz84u/O+VHA5woFkjcKIGPu+EAhZP8O/7L4kNjjCAqqNwjTk8r4Y11o/u5XlvhrcNsJy6Y3C7aPrvoupXD8yVd++AqQ1wvYtjsI/L+W+hgleP8Gg2b5KVTTC3HOOwrQS3r7KrV8/5GDTvrUiM8KGqY7CkszWviKIYT9qCM2+JecxwoXqjsIzl9C+xbZiP+Zsx74qrTDCQiqPwoJ2y76pLWQ/x/rCvmNwL8JQYY/CsGDEvkA7ZT8ebby+4TMuwuGUj8KZX76+PrdmPx4Wt7699SzCCtKPwhNVur5CPGg/zauzvkrXK8KdBZDC+m22vjByaT8fQ7C+r5Uqwvs7kMIPW7K+LuBqP1m/rL5PZSnCrW2QwmKsrr4dmms/OlypvpM3KMK7oJDCJJurvuDwbD9EyKa+YAgnwjXQkMJCcqm+cOVtP5z2pL4f8yXCaPuQwvTJpr41024/fqKivqK9JMLuMJHCdHSmvoW5bz8umaK+eawjwn9dkcJN2qW+TxxwPzYgor5UmSLCW4eRwrTBpL76K3E/S2Ghvkx1IcJmvpHCbFilvmzWcT8LL6K+lm4gwlHpkcKIwKa+1iJyP6yvo77cbB/CGxWSwqUqqL6cRXI/GSWlvm5mHsLrN5LCG2+qvgHNcT8zQae+EFkdwl1mksJVbqu+U7ZxP7o4qL4PWRzC8o6Swkaxrr6vDnI/cpqrvilfG8K+v5LCkbmxvnCVcT/lea6+4WQawhnnksJDPrW+SAlxP4vOsb6jdhnCKhaTwni8uL7ef3A/0By1vo6OGMI/OpPCUjK8vm9FcD9Af7i+bKcXwoZqk8LrgMC+M9FuPyZEvL5hzBbCZpGTwpVdxL555m0/yce/voj2FcIDvZPCXaDIvlK4bD+flMO+uhQVwvLtk8LHIM2++q9rP/6qx76KShTCARiUwj1Q0b7wK2o/YDrLvsNtE8IQRJTC9aTWvrPuaD+RB9C+QqUSwkV6lMJOcdu+BflmP4X4074v2hHCs6aUwlMJ375UYWU/Q9nWvtYKEcKez5TC7PLivrw/Yz9DyNm+XFoQwroDlcLRoui+aPJgP4Zd3r4ulA/CVDaVwryU677BwF4/+kDgvh/EDsIYZ5XCVorvvqP3XD/OT+O+RfYNwqialcLomvK+6ChbP0t15b7yMQ3C5NOVwszk9b7GcFk/Q9nnvp6hDMJdA5bC/mb4vhtjVz/rSum+AMULwkhBlsKojPu+7JBVP/B0674eIAvCBniWwjZj/b6J+1M/AXPsvk5QCsIqp5bCYVL/vjphUj/Cg+2+iq8JwgnjlsLbgAC/txVRP1F87r5iJq3CF3KaQQvWSj91MSQ/D8IvPwbZrMLCnptBHVlOP7OYJT+wwDM/JIuswvvBnEH/GE8/cXYnPwxINT8GRKzCQ+adQW98Tj/IYCk/H4A1P+/7q8JEDZ9BknhMP6myKz/+gDQ/arGrwjI3oEH2pkk/bRkuP/K7Mj9sZavCClWhQXm2Rj+tVjA/y8IwP4gZq8J+d6JBzt1DPyUIMz+xCy8/a9WqwvmVo0Eg40A/xXo1P2gULT8ijKrCkLCkQWGTPj8GfDc/EpUrP7tDqsKk/6VBNK89P8bTOD8HOSs/4PypwjQjp0EjjTw/htA5P2F7Kj9at6nCMTmoQQd1Oz+4xjo/TcQpP1d2qcIYValB5K86P8NKOz9FMyk/cTSpwvhiqkGzYDo/oEM7PwTiKD8H9qjCVHCrQcYqOj/cPTs/XaooP5qrqMLYg6xBJzc6PxygOj+Xeig/hGiowvyarUG/Yjo/8Lg6PzOvKD/JKKjCt5quQcOBOj+8tDo/UMwoPyrnp8JasK9BWrI6P3wtOj/RyCg/MqSnwtywsEF/iDo/afo5P+6LKD/WYafCGbuxQT3bOj/qgDk/WK8oP64gp8LetrJB5NI6P843OT8qiyg/Xt2mwpOks0G0/zo/IrY4P9eFKD83nqbCopq0QahpOz/gwDc/ZpAoPxdXpsLGlrVBodg7PwlLNz930Cg/5R2mwreCtkEwATw/f5E2PwuxKD9c2aXCPGO3QW2ZPD/dvjU/gvUoP0mcpcIjV7hB2EM9PyIGNT9sVSk/22ClwnJBuUGCHD4/IswzPwCwKT+cHaXCgiy6QcjOPj+sbTM/6TkqP/DipMJeAbtBso4/P9cpMj8Hdyo/iqSkwgnuu0FOI0A/p3gxP7zCKj8cZqTCFM28QTolQT+4kTA/r2MrP10qpMIZoL1BxYdBP7GsLz8YaSs/ZeyjwhKHvkFqm0I/4JMuPywGLD9Gs6PC+Fq/QRgwQz9psC0/DjwsP799o8LtPcBBDMBDPzoBLT/lgSw/20CjwnYnwUHhIEU/WucrP3tnLT/EB6PC0AfCQeqdRT/AyCo/uWwtP0/NosJI/cJBqKRGP5cIKj+THS4/2JaiwqrCw0E/lUc/7CQpP9mpLj85WqLCIp7EQfNkSD9xLCg/Aw0vP/knosIQiMVBb9xJP9noJj/f8i8/Z/ChwtRfxkGRXEs/6eolPwT9MD/Xt6HCYFvHQVatTD983yQ/u9IxPzmDocJSM8hBWl1OP6qyIz+R9TI/vEOhwpQ2yUFMzVA//wMiPzCZND/IE6HCGxnKQROnUj8w0yA/7N81PyjZoMJZDctBPzNVP91MHz+yrDc//KmgwtkIzEFqmVc//n4dP7AyOT/4baDCAQDNQTl0Wj+G1hs/HTg7P+I3oMIy7M1BjhtdP1dqGj8yJT0/ggSgwsX7zkFDjmA/I1wYP+uKPz/7zZ/C6urPQb29Yz+h5BU/vHlBP5Won8LW49BBrthmP0POEz+rfkM/wnCfwk/t0UH6uWo/+nwRP3YkRj+JNp/CQPHSQXY/bj8kLA8/42xIP90Ln8Ik9dNB9bBxP6CdDD8nfko/KOOewvIC1UGNTHU/JDkKPzTJTD83rJ7CExDWQTOZeT9+Bgc/EFBPP4GAnsKLGddBmcd8P1pJBD8u9lA/UVWewj4g2EH7PYA/in0BP2cRUz+8L57CvS3ZQfIlgj9j8/s+99pUP879ncItWNpB0guEPxaj9T7LylY/7N+dwnhS20EN0IU/gN7uPn5PWD/SuZ3CZXzcQYtUhz/k9+g+z49ZPzCRncJSht1BtOOIP48w4j6em1o/Mmadwq6V3kGYaYo/o/LbPv61Wz8ASZ3CMLDfQRv8iz9GT9Q+kXRcP6MdncI46eBB9G+NP55Kzj4cbV0/8P2cwvwS4kFanI4/8FrIPhjbXT911ZzC+2DjQQTljz+92MI+qptePz2tnMK1guRBKwKRP0HkvD5H3l4/RoqcwsvG5UEYCpI/0pm3PnooXz8SYJzCsBXnQWUvkz/Bo7E+NW5fP/c/nMKuM+hBq7yTP9imrT50MF8/GCKcwqZj6UGHgpQ/VjCpPho0Xz++/pvCvaLqQWYklT+CMKU+KxdfP5DXm8K4/etBDaCVP+VKoj6vDF8/1L95wv7EfcI2ChRAJM4xv8K7Pj+SzHnCtKN9wmD4E0D2QjC/5VU+P+DHecK0g33C9CAUQJaHLr/Q7zw/pMR5wgRofcK9HRRAqFAtvzV0PD8GxXnCFVB9wlbnE0DdKSy/Bcw8PyjGecLTQn3C27gTQOSXK786RT0/tsd5wuY4fcIYnRNAz5orv3m1PT8gyHnC9S59wn9kE0AWkCu/JpM+P8TMecK8K33Cx5ATQPTuK78YDD4/BNB5wjAqfcLA1BNAR/0sv6dzPT/K2HnCHC99wi0qFEAYGy6/P5s8PxDbecLYKn3C95UUQFa4L7+enzs/eNt5wnQtfcKR9BRAjJEwv2iCOj8I2nnCpjF9wr6MFUBCuzG/eqA4PwLnecJWLn3C7PcVQMioMr9DWDc/quh5woA0fcL3kBZABHIzvxRJNT8I53nCvjR9wvH2FkBhvjO/EtIzP8TmecIONH3CcIAXQK7/M7/eyTE/1PB5woovfcLR4BdAMj00v09kMD+u7nnCBjJ9wiotGEBxgDS/31AvP8rsecIyMX3CIY8YQNaGNL/mzy0/uO15wjA2fcKY5RhAWtI0v3uYLD926HnC1Dt9wgcvGUAC9DS/N4QrP6TqecKDPn3COm8ZQBI0Nb9+oCo/Nu55wmg7fcKKyxlAJqU1v1BhKT+Q5HnCwEJ9wmz2GUDFmDW/C7QoP77kecIHRX3CIhcaQMBnNb/PICg/fOJ5wqdRfcKjLRpAOSk1v5OwJz/m5HnCy099wtg9GkCNrDS/WEEnP77qecLiUX3C1jMaQAxiNL/xSyc/eOJ5wg9WfcLxERpApCUzv9hWJz8A4XnCZVV9wif8GUB7kDK/a3InP87pecKVV33C2cMZQGLFMb9z/yc/7uF5wk1ZfcKiixlA/Esxv2urKD8U5HnCVlt9wmJCGUBzezC/fncpP1LrecKUXH3CR+UYQHg5ML+zyCo/ivF5wsZgfcLSaRhAP4Yvv8xjLD+6AXrCsV99wjgfGECEhS+/aYctP8b/ecJ4Xn3C5rMXQKK6L7/XQS8/jAZ6wkRbfcKISBdApbovvznoMD/eCHrCTFd9wv3hFkB/zi+/4oQyP2ANesJPVH3CiVwWQJnGL7+gkTQ/pBR6wgtffcIP/hVAZY8vvz7xNT90GXrCfVh9wvqFFUBmti+/at83PxAkesJrUH3C8hUVQHrCL7+nozk/YCd6wjtPfcLKyRRAroQvv5m5Oj/uN3rCUE59wkRgFEAAfC+/wlw8P546esKuR33CaBYUQHRPL7/ocT0/1D16wupFfcJ92hNA7BIvvztIPj9wT3rCHj99wheVE0AtDi+/pF0/P+JTesKzP33CiF8TQImFLr/I+D8/rFF6wtsxfcItIBNA+Bwuv5vJQD9kYnrC5Sd9wvn5EkB3ti2/wjVBP/BqesLvJn3CieMSQPI3Lb9aV0E/eHF6wvMofcKG9xJA8AUtv0DwQD/GhnrCPh99wu7OEkAkoCy//WVBP3iOesLsFH3CNd0SQI73K7954EA/JJR6wu8OfcKU6BJABGErv/VuQD+MrHrC4gR9wv0dE0A6fyu/LqY/PyiwesJS+XzCuzMTQPg0K7/ELT8/VMN6whX9fMLnYBNA4gkrv7ZlPj9y1HrCcvN8wu2fE0APFiu/RW89P2jsesI68XzCluMTQOo1K79kbzw/7P96wj7mfMK3DRRA7jorv7nJOz+QJXvCguF8wsVnFEBtTSu/k2s6PxA0e8Ku0XzC0q4UQNplK78bXDk/DFJ7wn3LfMIfExVAoH4rv2HZNz/IbXvCqsF8wiRlFUBV1Su/K7o2P8Bxe8IatnzCz7gVQKB7K7/TSTU/AJR7wn+mfMIIHRZAgt4rvwDpMz/YtHvCKJF8wgt9FkAWfyu//EcyPyzOe8IPhnzC8+YWQCVsK797oTA/1OF7wit+fMIaNBdAHM4qv0czLz+GD3zCZGJ8wtmKF0BlTSq/Oa0tP1IYfMKqW3zCXNgXQHWJKb85MSw/rkZ8wqNKfMLrSRhARC8pv1xWKj/6VXzCtDl8wvSRGEBBPyi/Y+EoP/SFfMIGP3zCVPoYQGyxJ78xGSc/8q58wskmfMKCShlAmicnv9awJT/iwnzC2Bt8wo2oGUBhSSa/ZvQjP2omX8K92YrCUegXQJ0oPL/ReTM/wjJfwrnJisL88xdA17w6v+G3Mj+OLV/CNLqKwpZAGED5Nzm/IOgwPyopX8LMrIrC42IYQHxBOL9p/C8/KClfwkShisLdShhA8lo3v+v/Lz/mKV/CT5uKwr0xGEDGCDe/DEMwP9oqX8KjlorCLCUYQNosN7+agzA/gCpfwuuRisJy7hdAICc3v4dbMT+oLl/CsJCKwnUOGEBtYze/IfQwPwIxX8L3j4rC6zkYQJEoOL/tlTA/IjlfwvKSisKDbxhAlN44v8cIMD8sOl/CIJGKwtC1GECXDDq/IGgvP1g6X8J3korC/esYQBKDOr9/vi4/CDdfwoqUisJ7UxlA00g7v8tuLT/uQV/CAJOKwm2TGUCQ5Du/OqwsP/JCX8IYlorCw/0ZQKRnPL+zNis/jD9fwgqWisIdORpAuH88v25TKj/iPV/CPZWKwiecGkD4pDy/m9coP1hGX8K6korCEtwaQGzPPL/D6Sc/lEJfwleTisKNDBtAmAw9v2FAJz9QP1/CQpKKwvZYG0CnDz2/lxImP0g/X8InlIrCWpsbQEtXPb9gJiU/wjdfwhuWisKhzxtA5GU9v/xcJD+kN1/CFJeKwlP+G0ChkT2/i7QjP2A5X8LplIrCxEocQBLmPb+WpSI/bi9fwv+XisIrbBxA5MA9v0UUIj+qLV/CipiKwj+EHEA6dj2/R5ohP9YoX8IqnorCGJYcQIkfPb+dNCE/GCtfwjOdisK0ohxAGo88v8bOID/UL1/C+Z2KwgadHEA+NDy/M8QgP64mX8KCn4rC+4IcQET0Or9XtiA/PCVfwrmfisIodxxAlFw6v5CtID8ILl/CRqCKwmVJHEAPlTm/4RchP9IkX8KgoYrClh4cQBcdOb9bkyE/ZCdfwp6iisJq5RtA5FQ4vzopIj+8LV/CwKOKwvuYG0D7Fji/ID0jP5ozX8L4pYrCmC8bQENmN7+clyQ/MENfwsGlisJo9hpA/2g3v6Z4JT8qQ1/C+6WKwumfGkDPmDe/LN4mP1BKX8JypIrCIkUaQI2UN7+cQSg/IkxfwhCjisIj9hlA0J83v0d9KT/eUV/CHaKKwsR/GUCLfDe/KkMrPwBZX8Lbp4rCKzYZQHMrN7+vRiw/il9fwsakisKczxhAD0U3v1fnLT+kaV/CbKGKwtByGEB2Pje/flUvP/JsX8LFoYrC/ToYQLPsNr/yEjA/3H5fwkGhisJS5RdAr942v6piMT/ag1/CMZ6KwrCuF0B+pja/6yUyP9CIX8JFnYrC1X0XQHBVNr8FyDI/eppfwqeaisI8SRdAqVE2v7GYMz8on1/CUJqKwlkfF0BkwDW/jQQ0PzKfX8LZkorCqe0WQKZdNb+QojQ/rLBfwkaNisKPzRZAL/I0v4D2ND9wul/C0IyKwhm/FkDzeTS/fP40P1jEX8LJjIrCTdgWQNZSNL+diTQ/lNlfwnyHisIZrRZAkPYzv98PNT8S5F/C8IGKwtC/FkADXzO/UoY0P+jqX8JFforCW8oWQKfPMr/pIDQ/TARgwil4isKJ+xZAxRIzv/t4Mz/uCGDCoHGKwpkGF0C81jK/PzQzP0wfYMJEcorC+i8XQM7OMr+fjDI/sjJgwsJsisJPZBdA4fwyv9nPMT+iS2DCmGqKwhGeF0DHSTO/NgoxPxhiYMKaZIrC3LoXQIh5M7+FqzA/bohgws5hisIKCBhAkLozvzqULz9WmmDCZleKwlpBGECBBDS/Pc8uP3q7YMKBU4rCiJYYQF5WNL9hny0/SNdgwhJNisIM1RhAFOQ0v8ngLD9a3WDCeUaKwrseGUB1zTS/RbUrP4IDYcJzPYrCaHAZQLZnNb//ryo/rCVhwpYxisLiwhlAsT81vw5cKT/wQGHC4yqKwkYbGkDFZTW/zg8oP/hWYcIrJorCil4aQMUSNb+G6CY/3IVhwvoWisI1pRpA6sE0v5O1JT+YkmHCiBKKwvrnGkCJMjS/F3skP5rBYcJQCYrCeE0bQKQRNL825CI/8tNhwvT/icKijhtAeUkzvxydIT+ABGLCsgCKwi7oG0BW3DK/xhogP2YyYsIl84nC5zEcQOpyMr8N2B4/YkpiwojsicIGhxxAoqoxvzNIHT/Qeo/Bh8i1wjTPJ77rK3c/MUklvop2j8F+vrXCgGsivsX5dj985B++arCPwUm2tcKpYxi+n9B2P6zuFb4T6Y/BsbG1wkZdDb4MoXY/c/8KviQfkMGWrbXCV0wAvvgMdz8Pcfy9NECQweyptcLKBOy9GJh3P61m6L2NapDBLKS1wilk1r0vYXg/uGTTvSKEkMFsobXCL6zHvesPeT+EH8W9EMSQwV6ftcKAV7+9x5V5P/4Uvb2A55DBcpy1whp2wL3Kl3k/YjG+vZIWkcEfnrXC2s7AvV9yeT8de769WkaRwY6atcLV88S9Te54Pwxiwr1Xj5HB9Jq1wrT4zb0E2Xg//UPLvVG3kcF6nLXCIIzZvVUueD8/bda9nOWRwZCatcIBFuK9LRp4PynT3r08IpLBUp21wtIC8b1HuXc/o2PtvXJSksFfoLXCOSEBvhopdz8/I/69im2SwcqdtcLITQ2+0rx2P9j3Cr7UlpLBup+1wpvGE77YbnY/SEQRvpDAksFZn7XCKBQcvg4Odj8UWBm+ot6SwWuctcIDhCG+sWh2P8rRHr5IC5PBz561wsSJKb7xXXY/uromvhkWk8G8n7XC/V8uvtojdz/TxSu+BjGTwRyjtcL4zzK+0eR3Pw5uML4YS5PBkKO1wl+YNL7K4Hg/0Isyvkdsk8FaprXCCqQ2vruYeT8S1jS+9YCTwSSotcJzoza+MJ96P1szNb4zepPBrKy1wmmtN74QDns/hWQ2vgqUk8GSrLXCAM82vp7Xez+lzjW+LbOTwW2wtcLJMzW+DJd7P54cNL61xZPB/LC1wjlEMb42OHw/+2Uwvgjnk8GftLXC6f8uvnrsez+5By6+kQOUwWi3tcJ12iy+QQ98P3buK76SI5TBjbi1wv7xKL4/fHs/ItYnvpQdlMHdwbXCJsQnvgZ8ez/ZqCa+gUCUwTzItcJ3mCC+vXV7P6J/H74IVZTB3sq1wshYG76Wons/i1EavoVtlMGvzLXCv+kWvi8bfD96CRa+FIWUwebTtcLMahG+mJV7P+NnEL7wm5TBc9W1wh9iC75673s/nX0KvsitlMHP2rXC1iUHvuTXez/aPga+7auUwc7ftcLHYQK+Pa98P4O2Ab6UyJTBw+G1wkRC+72QT30/zED6vZzMlMHD6bXCYcfxvevPfT+TBvG9ftSUwSTstcJroem9hA5+Pxz/6L015pTBQ/a1wtBn5L1aoX4/6AfkvQ3XlMEu+7XCs7zdvdxCfz85ot29lPuUwar8tcIei9O9nPJ/P8i1070u5pTBmAS2wqy50L0BT4A/cSjRvSfvlMG/B7bC//rGvQyEgD/Sice9MP+UwawKtsLT/MK9o+uAPy7Ww72j7pTBsw22wuy7wr3PKoE//MTDvY7ilMGiDrbC6TXAvS6JgT98gcG9Lv2UwYYQtsIS0729OKSBP7Utv72aDZXB8BW2wkRju73j14E/at68vakqlcGeGbbCYn+8vTUGgj9QH769EgSVwY4btsIddry9ST+CPwBAvr0WG5XByh22wgzmu71kpII/n/i9vWsUlcFNHrbC1xi/vSC+gj/dSMG9FweVwYojtsIllcG9QeuCP47vw72+NpXBex22wqlkwr2LEIM/P97EvTYelcGbIrbCJzXJvaUUgz/5ysu9KkqVwYIitsIJJs69NjmDP7jr0L0eRpXBMB+2wn530b0FOYM/7knUvcpvlcF0HbbC5W/YvVEugz+iVNu9R4OVwYIatsJJTNy9lT6DPypO372RnZXBUhe2wu66473KWYM/7fLmvejJlcHPG7bCPCTuvSz9gj+eMfG9z9KVwQgZtsJIRvC9uOuCP4dL871m3pXB/BC2wiWJ9b1KooI/DFz4vdgZlsEWDbbCDmz9ve1kgj8/DwC+nGeWwa4LtsIMUwK+rOuBP8t6A76cbZbBTQa2wgqEBb4sj4E/B4UGvgKtlsGk/LXC/7QHvqgGgT9fcwi+ffaWwWL5tcKK0Qq+jFKAP+U0C77bKZfBpvK1wqmnDb7oYn8/ErYNvn5el8Hv8LXCD64Rvk8Afj9AWxG+J7iXwaDutcJCbhO+LZR8P/WyEr6FNpjBTOK1wtKoF77M1no/4GcWvs5XmMHk2rXCi1sZvpyWeT+HuBe+uLRZQhvPmMKUvsk+K6lZPykXvT5/I1tC5HqYwmgazz6x+Fs/zjLDPhSRXEISIJjCuJ3UPr37XD+k+Mg+wvldQi7Dl8ILkdk+GUNdP7/pzT7OZl9CJ2KXwgkw2z6ayF0/xLjPPibYYEK9AJfCearZPkRkXj9ff84+Kk9iQjiglsJCR9c+lcFfPw6/zD6czmNChkCWwj7Q0z5dT2E/oQHKPi1NZULD55XCvlTOPt4oYz8hY8U+2NBmQiOMlcK5psg+6rBkPwFrwD6YWWhCQziVwtEsxD4XwmU/am68PhThaUIK4JTCLn++PoLeZj/8Q7c+2GVrQuSKlMKWSbg++o5nP9losT6h9GxC+DiUws+rsj5sHmg/WhasPsN+bkKE6ZPCe4CsPljHaD/9QKY+t/xvQoeXk8L4p6c+DDJpP7aioT72hHFC2UqTwrIdoz7DqWk/MVWdPrcFc0JE+pLC3yuePrxbaj+4s5g+bIR0QpKuksLULZo+MBZrP1MClT76/XVCxmKSwoZNlj59ems/21KRPttpd0KdFZLC85+SPpscbD+A5o0+TOR4QkbMkcJFnJA++JBsP8cNjD4iU3pChoSRwrmijj4y2mw/QjKKPtm3e0IOQZHCY7qNPnfubD/kU4k+fhV9Qsj4kMJhF44+P9NsP4+niT56i35CQrOQwnBujz5jzGw/VvaKPhvZf0JNbZDCdhWRPtWzbD+ujow+xKWAQjMmkMJZfZM+qXZsP9/Zjj4QVIFCAeCPwkwllj67+Gs/sFCRPtr4gUK3l4/CDMKZPlB0az8OtZQ+7KiCQm5Uj8KH450+mExrP6G3mD78UINCqxGPwkcCoT4SuGo/4JmbPk7xg0Jfyo7CxtWkPug5aj/cM58+Qp2EQvOEjsLTPak+bdFpP1hmoz7oQIVCKUCOwl7trD7WGGk/vcemPiLlhUJP/I3Ch8KwPvvSaD+EdKo+IISGQhW2jcIGt7Q+f8loP9xVrj6OIYdCE3KNwnBNuD4Slmg/YMyxPiDFh0IgKY3C+6K8PtyMaD+mD7Y+U2eIQsnijMLg0r8+wlhoP+whuT5yB4lC4pmMwsBuwz4fQmg/paq8PtCbiUIWWYzC++bFPqeraD+qRb8+LjqKQgoKjMKvTsk+KZJoP+2bwj6g3opCeMuLwuw/yz7Jemg/A4DEPmNxi0L/h4vCeQDOPvtXaD+TLcc+pgaMQipBi8IjrM8+P5ZoPzLwyD5YpIxC/vqKwrbT0T48f2g/egvLPtg+jULOs4rCJEjUPqF0aD/OeM0+ltyNQtpuisJbadU+7tJoP+HAzj54bI5CLCaKwg1h1z414Gg/KL3QPh4Gj0J46InCwaXXPvnpaD/tBdE+0qGPQpClicI5cdg+LDZpP9Xx0T6AMZBCzmKJwu572D7vdWk/5xfSPovDkEI1IonCCzzZPi5faT8/ztI+nFGRQhDjiMIPJ9g+W+tpP1z10T7p15FClqWIwhAf1z5G5mk/CuvQPvJukkLZZ4jCQLDWPpy7aT8LatA+j/ySQggwiMKSd9U+kOtpP8lFzz5MgJNCLvOHwukA0j4JbGo/egXMPk0BlEJQv4fCfWzQPnZ8aj8/eMo+AomUQq2Ah8LnH80+GNFqP5xPxz5sD5VCHE6Hwt8PyT4v02o/KkPDPveQlUL/GYfCjULGPnI4az9moMA+GheWQsjphsJXJsI+WEdrP5+OvD45kJZCOb6GwiPNvT4pWWs/tkK4PrsGl0IylobC2ni5Ps5waz9D/7M+8HyXQipnhsKxgLQ+AZtrP2Ahrz6CAphC5D6GwjFlrz4W+ms/KTSqPqZymEL+IYbCkuWrPijvaz+huqY+y+2YQp30hcLw0ac+HqZrP6Oboj63YplCy9KFwnWhoj6iuGs/F4SdPvHTmUJCrYXCfyifPrzFaz8EHZo+AUyaQueLhcL2qpo+oEhsP2balT4FwppCIW+FwsxElz6G92s/p2qSPjczm0KvUYXCdMeTPkcVbD9YBo8+brWbQlY3hcK2DJA+OkdsPwBsiz7fMpxCrBmFwp+pjD5CUmw/Ch2IPjaanELMAoXC7LOKPh+ZbD/ZRIY+o/+cQhLkhMK7d4g+Hj5tP4VAhD5Kep1C6s2EwskDhz5Hbm0/m+CCPi+1rMLW8BTC650qv/myQz9tWhy/imyswoZRFcLw7Cq/poFEPzDwHL8YJazCta4VwgzWKr838kU/R1kdvynkq8L3CRbC48Ypv/kwSD8QEh2/YaKrwthmFsLyqie/tk5LP7QHHL84X6vCAscWwnxFJb/iTE4/HqMav+AZq8JSJxfCULUjvw64UD/I3hm/EduqwnqFF8IwByO/3GRSP9W8Gb+Mm6rC/+QXwuKnI79nAVM/AJMav6ZcqsKrRBjCaHclv5SEUj/aPhy/0BqqwiOvGMJrHSe/o5RRP8CZHb/s4anCRwoZwgjMKL/dE1A/Bssev2GgqcLKbBnCcuwpv8w7Tz/OpB+/PmKpwobWGcIJ2yy/SLpNP7wWIr9hKKnCfi8awh8VLr++10w/BQUjv/XoqMIVoBrCmaUwv2w0Sz9VCCW/6K6owmsCG8KwYTK/HV9JPykgJr9baqjCPmUbwnrjNL/MPkc/i+EnvwkwqMIZxxvCRfM1v6x6RT98TCi/R/anwkg3HMK1cDe/AK1DP2ggKb8MuKfCe5EcwtVkOL8ZiUI/N6gpv718p8L0/hzCIcg5vw0vQT8tiiq/cD6nwgVrHcKkQjq/muBAP1PnKr+8AqfCgtIdwv+VOr9u9kA/WEMrv+jEpsL2MR7CSjQ6v+YaQT8b7yq/KYamwu6VHsKihTm/mcJBP2B/Kr8XTKbCHPUewq3sOL+eX0I/FSEqv38IpsJ7XB/CY104vw94Qz+0+im/tsulwgS3H8JN5je/gkxEP57SKb/1lqXC2BYgwsNuNr/SZEU/18Aov5JPpcJveCDCHJs0v44nRz/ojye/5A6lwmPOIMKokjK/u7hIPy4VJr+t0aTCCichwhcPML/6C0o/0QUkv3WPpMISeyHCI04uv7huSz/iviK/nFakwibWIcJDdiy/d+BMP2lkIb9TE6TC8TMiwj1uKb9fcU4/jt8ev9zWo8KsgyLCYMImv97HTz/ioRy/dJujwnnYIsJpKiS/0BlRPzh1Gr+XVqPC4TAjwtz/Ib93mVI/TMQYvz0Zo8LnfiPCatIevxWZUz8c4xW/MdeiwlvJI8L05Bu/CuhUP/NaE78+jqLC9SQkwnBKGb+hSFY/JioRvxNPosKacSTCEtsWv6uNVz+SGg+/IQaiwh69JMJJ8BO//yhZPwCnDL/+x6HCKwIlwttGEb8cF1o/GkAKv0p+ocLlSCXCS6cOv9s+Wz/98ge/qDehwtCQJcKbRQy/VT5cP13XBb/d86DChtIlwuenCb9TV10/nIUDvw+loMIFHCbCmRQIv58wXj+PLAK/RWagwoZaJsJ+Cwa/ZJVeP1w9AL9JIaDCc5kmwniuA79SuV8/fln8vgLRn8JW3ybCjU4Cv0iIYD+VBPq+dJKfwtodJ8LSCwG/CepgP46w976WSZ/Cc2Enwv70/r4I7WE/+RD1vmkCn8JvlyfC73/9vu/AYT/5hPO+FLuewszYJ8Jv2Pq+v2JiPx4u8b5ib57CgQwowliy+b7RPmM/T3XwvtUmnsL9SSjCWaj4vhSsYz/poO++8eOdwqaCKMKlvfe+Vd5jP3HO7r6klJ3C474ownOo977YLmQ/6uDuvoJTncLC7SjC7EL2vjTTZD/4yu2+AA+dwmAyKcKxCfe+Y3VkP2Jk7r7aypzCFGMpwuGt9r5R2GQ/3TjuvqKEnMI3kynC79P2vnqyZD95TO6+eT6cwjvLKcLM9ve+hadkPyJr775L/5vCNQsqwh+8+L6cY2Q/sg/wvh25m8IARyrCbwb6vmpiZD/kWvG+HHibwsaDKsLqA/u+icxjP6gO8r7uK5vC5L0qwitr+74hkmM/EVnyvpzlmsII7CrChJ/8vp6AYj/qBPO+qJ6aws4uK8K8PP6+DMthPyFH9L4oXJrCPmUrwqLG/74jl2A/qDT1vjAOmsJolivC8jAAv8ffXz8ecvW+rsuZwuDWK8LeyQC/EedeP8cj9r5hh5nCkxUswkZnAb/r1V0/sdD2vslamcLeSSzCb84Bv2ueXD+2/Pa+hgKZwi2ULMLwXwK/adZbP8K1976CxJjC4d4swlAXA78Itlo/2Ir4vgx0mMKiAy3CjEEDv6q2WT+pWPi+TTOYwjJaLcJ9+wO/uydZPyZ9+b4NERbC9Z2twm6XGkBVsD6/M7EpP7wWGMKE06zCrMgZQHcUP7+gEy0/0uwZwlIBrMKhZhlAH4I+v69kLj8jvRvC5y+rwsDYGEAlgj2/1DowP4GVHcLwWqrCUjwYQMqFO7/E5TE/h1ofwhyNqcK7sxdA7505vxRGMz92IiHCVr+owtNSF0DGLje/3swzPxXiIsKO8afCaKcWQJp2NL/NWzU/Yp0kwrgnp8JvWBZAnc8xv9h7NT9hTibC0FymwuDYFUDXXi+/a3A2Pxj4J8KIoKXC7GwVQAItLb8cLjc/lKEpwvzTpMIPFhVAW0wrvyW4Nz8oRCvCtgqkwlCzFEBecim/vXE4PzzVLMIPQqPC458UQPbNJ7/qBzg/znsuwmt6osJUjBRAbNEmvw/nNz9kEDDCQ7ChwpuFFECj6iW/CZ03P12lMcKR6qDCo3AUQKZrJb/mtzc/MiYzwrYloMIvehRAYhklv6tuNz+ApDTCClifwu6IFEBKAiW/6io3P+4bNsKNlp7CGnYUQF7dJL99ZDc/Hos3wuXNncIFlhRAtOwkv1buNj/G4jjCZA6dwuWcFEBtxyS/RMM2P0RhOsIeTpzCSKgUQFh1JL8sczY/YL07wjSQm8LBmxRAy9Ajv5pcNj8FLT3CqM+awiPkFEDeTCO/tgk1PwtxPsIaFprCV+QUQHJgIr/ZojQ/qNc/wnhbmcLA6hRA3/Ihv7VaND9PJ0HCzaeYwqkAFUAWNiG/jbQzPzZ2QsL555fCOxAVQOOhIL+kODM/rstDwr4sl8KfHxVABNkgv/gUMz8ZDkXC3HSWwsMXFUCUEiC//d0yP4hnRsL9wJXCPQwVQCYJIL9TBjM/eqhHwsoGlcJE/xRAn8UfvzcbMz9q80jC+FKUwrcXFUA5vx+/VboyP9IsSsKDopPCGNwUQJx5H7/ggTM/todLwmDjksJWthRA7g4fvyXlMz+v0kzCazWSwlKTFECdFR6/tf8zP2UJTsLdfZHCYFwUQNjcHb9qujQ/YzJPwl7HkMIsDhRA4Mccvw9uNT/me1DCCBCQwoS+E0BdLBy/GFw2Py2sUcICXI/CRWsTQCXoGr+MDDc/hsdSwmSkjsLa5BJAq20Zvy1oOD//9FPCtfSNwm+yEkABeBi/Mbs4P8FFVcLDP43C42ESQBiYF785izk/2GpWwqqIjMISGhJAkAEXv6daOj9xl1fCJ9WLwou8EUC25BW/dD87PxzGWMLsHovCrnIRQANLFb9/FDw/1ORZwqZtisIbMxFAFpcUv3K1PD/QEVvCCbeJwjzwEEC0NRS/Y4k9P94lXMKyB4nCN8IQQN6VE7+O7z0/SEJdwrdUiMKdcBBAKNASv4zMPj+oZ17CFZGHwjruD0CacxK/mZdAP+qAX8K24IbC9bkPQGeoEb8VAEE/VJ9gwhInhsJYbg9AcFARv/L5QT/gtmHC/3GFwrUYD0D3thC/6fpCP4zsYsJ7s4TCspoOQOrvD7/pgUQ/tvNjwrUBhML0Nw5AOWAPv1i6RT84AmXC9EuDwonSDUBNCQ6/2plGPzoyZsIgioLCf2QNQOvFDb8MJEg/bi9nwu/OgcIc2wxAiG8MvyOOST90OmjCDBGBwp50DECV5wu/mdhKP4xJacKsY4DCwxsMQAG9Cr8Pmks/ZGZqwjdVf8L0oAtA9RsKv6QmTT+4ZGvCUNB9wnIzC0AMcAm/s3lOP8JwbMKMe3zCTeoKQK6oCL+WL08/1JRtwmH+esIXdwpA8UgHvxQ5UD+snG7CKpJ5wrkrCkDuCAe/1D5RP8Kpb8KYIHjCzuUJQOq0Br+0JFI/TIlwwn2ydsLsgwlABs4Fv3UpUz9C23HCSU11wsiACUDCawa/TItTP3TXcsJS0HPCh1AJQBSLBr/fWlQ/OBB0wiBWcsKuIglAwTEGv4bfVD+SxHTCygBxwpAQCUA+zQW/UPBUP8zsdcJqkG/CU/4IQCJTBr+ugVU/kNt2wqYhbsIN5whAwrwGv80XVj920nfC17hswsQKCUBIOwe/SM9VPyABecLZTWvCCSMJQD3lB7+yy1U/Dgt6wsT8acIQPwlA4L0Hv/dGVT+uFnvCRolowsdnCUBElgi/ABtVP3Tze8I/GWfCeW4JQGw5CL8LzlQ/Yn6+wnnROsDm+RTA/8szvznRO7+LAL/CSglMwPQgFMAjYjG/aTA+vyuHv8LOV1zAxYITwNwSML9rG0C/ghDAwq4xbMAQBxPAE2gvv7LDQb+wmsDCq+B7wDYCE8CMpS6/8H9Bv8ogwcJt3oXAL00TwC6jLr8EUEC/FqPBwmmujcAcnhPAT7suv4YUP78dKsLCVHSVwGUWFMAQ9y6/Iks9v76uwsKOQJ3AWGUUwC34L7+Qfjy/BzbDwunepMASxhTAL+kwv81iO782sMPCKiSrwMhjFcCzMDK/rHY5v/ozxMIeg7LAJNsVwJWQM79yLTi/zK/EwgwFusC3ZhbAbCY0vzE9Nr/uMsXCsj3BwHTQFsBpLTW/dQM1v8qzxcLDfsjA6TQXwGgONb9NZTO/+SvGwobmz8BwcBfAwvg0vyNvMr/ioMbCMaPWwC7dF8C6wjS/wKgwv1YXx8IojN3AHhkYwKCLM7/HPS+/CJLHwt6/5MA1ShjAzPUyv8k/Lr9JBsjCSYXrwLNsGMDrcTK/KIMtv+p7yML/XvLAynwYwBQbMr84IS2/6OjIwhIA+cBAWxjA/PcxvxOXLb8kUsnCfJr/wB1qGMAWKjK/pHAtv6rCycK7RwPBFUgYwPjdMr+fPi6/sDPKwppZBsGgTRjAJO4zv8WVLr+5msrCRXgJwXI5GMAKFjW/QlwvvyAKy8LgswzBehIYwCiZNr+qkjC/YWnLwgW3D8Ez3hfAcdw3v7TlMb+51cvCPcwSwfinF8D8CDm/1Tgzvyk4zMJ/tRXBdpkXwK6POr8NEjS/lKLMwqq+GMGCUhfAwkU7v8V6Nb9g+czCoLsbwd8WF8CX7zu//bE2v5NezcLEqR7BafIWwHIDPb/Ptze/38XNwvV8IcGlxxbAtO09v/7GOL/1LM7CekQkwYaiFsCBqj6/Fa05vwaCzsKU7SbBBLYWwFFvP78Arzm/ne3Owg7pKcEnmBbAtPo/v2ljOr/yWM/CZIkswYK/FsBMx0C/wBY6vyC8z8JDMS/B+s0WwLKbQb9EMzq/oBfQwhHXMcG2DRfAd6JCv3KZOb+7edDCd3g0wfRhF8DZkUO/yqA4vy7W0MIeLzfBh6gXwIp+RL+f3je/+jLRwsCxOcFECRjA6a9Fv7DLNr/Sk9HCcAI8wTSIGMDPmka/Ih81v/j20cKFYD7BPi4ZwOCUSL/bOzO/q07SwsnYQMEyyRnAothJv3U8Mb+9rNLCRTJDwT1hGsDOOUu/ElMvvzsP08IrjEXBJS4bwDDsTL+BrSy/1WrTwlwWSMFH0RvACn5OvyOlKr+9ydPCfD5KwadrHMDX0E+/sqcovzci1ML5ZkzBVxwdwEr2UL8hPSa/y3jUwuTkTsECwB3AnNhRv2/vI7/j2NTCiTpRwQRPHsCprFK/AfEhv9ks1cKdIFPBsMIewNpSU78SUiC/bIXVwjBdVcEYQB/AaKFUv1nEHr/I6dXCIHxXwV/DH8CqTFW/iuYcvzYz1sKo1FnBSwMgwEmXVb/v+hu/9IzWwsunW8GoZCDAZShWv2+eGr9W69bCs9ddwUyQIMDWfla/Lgkav/pB18LM4V/BvsYgwD7+Vr8yVRm/8IzXwlsgYsEk5yDARUVXv3HoGL/X6tfCA1hkwcn2IMA90le/ctYYv85K2MIiXGbBQPcgwHnkV7912hi/7JTYwqhFaMGbBiHA2g5Yv6KpGL/g6djCroxqwR/xIMAmLli/TgsZvwdH2cLQcWzBvt0gwMsjWL/3Vhm/gorZwqJQbsHzrSDAHqxXv97yGb/P19nCx0JwwaKZIMCXG1e/dhYav/0m2sJWUnLBZl8gwLgVVr+xrRq/E3HawhaOdMHROCDAjkRVv7IFG78Kw9rCgEJ2waIUIMALe1S/EVYbv3wX28KdRXjBx+UfwI98U7+3vxu/F1LbwnhOesE7sB/A8MBRv4AEHL8NntvCvhl8wdeVH8AHiVC/kwYcv7Tu28Lg133B/6IfwMZrT7/DcRu/IS7cwivTf8FGhB/AZHNOv/iZG7/sa9zCKsaAwfxsH8Au2ky/p20bv9W33ML3zYHB8G8fwBQGTL9NGhu/2vzcwgCigsFvXB/AuuFKv5sFG79ANd3CfpSDwa5SH8BPUUq/v/sav4s2wMKxunfAKSMVwAB3NL8tdDu/abbAwvdmhMCHQhTAx/Exv8PnPb8CO8HCTHmMwEyZE8AwijC/VfU/v6TCwcKAUZTACBQTwDvML78dvEG/REvCwmEVnMBzCBPAMwAvv2yPQb8X0MLC2e+jwPNOE8CJ+y6/WnBAv4pRw8LrqqvAg50TwLYaL79DQT+/fNfDwvBas8A2FRTAp18vv9R9Pb+FW8TCEhG7wDRnFMCNZzC/rqc8vxPixML0lcLAoMsUwPJlMb94gju/rVvFwim8yMDtcBXAPr0yv+x9Ob/63sXCDv7PwHLxFcAyMDS/YRc4v61axsKTYdfAGIcWwELcNL9rBza/c93Gwlp83sCR/RbAAvs1v9qjNL9WXsfCAaHlwDVwF8Bp+zW/sNkyv0TWx8K95uzADboXwH3+Nb9ktDG/BkvIwqyB88A2NBjATeQ1v83DL788wcjCnUr6wLV8GMAiwzS/wzAuv/A7ycL/rADBybkYwHg+NL/BCi2/+a/JwqT9A8FG5xjAUs0zv7kqLL96JcrC5VcHwav/GMBnizO/wLArvx+SysK3lArBoeYYwFZ1M79/Ciy/p/vKwrrQDcGY/BjAvrgzv7XOK7/Aa8vCzTgRwXHgGMDHfTS/X4ssv8Dcy8KhNxTBwugYwA6jNb9v3iy/NkPMwvlCF8Ew1xjA+t02v3ygLb9esszCWWwawUKxGMCycTi/N9cuv5IRzcKTWx3Bi34YwBTHOb95KTC/rH3NwgpcIMGbSBjAsgQ7v6SAMb/L383CYzAjwVg8GMDWnDy/nFUyvzJKzsIMJCbB4PQXwLRgPb8bxTO/oKDOwkoOKcGiuRfAgB8+v2cCNb/2Bc/CVuUrwf2UF8AKRD+/Sg42v21sz8Kfoy7BrGsXwMw6QL96Gze/UdPPwv5VMcEtRxfAiwRBv1kDOL+ZKNDCmuUzwUFfF8AN3EG/3vg3vwmU0MI7zDbBBEEXwB11Qr+4szi/Lv/QwsFTOcGbbBfATk9Dv8FZOL9tYtHCkuE7wUSAF8DRMkS/qmU4v9290cJHbj7BMccXwHNHRb/JsTe/bCDSwhD0QMGXIhjAgkRGvyCfNr/ifNLCFJJDwepyGMByOUe/FbY1v+rZ0sIr+0XBzt4YwNyCSL8RfDS/GTvTwtMrSMGoahnA3HhJv1OcMr/LndPCoGtKwSogGsCihEu/gHswv5T208LGxUzBg8oawG/aTL9hQC6/QFXUwmf/TsEHdxvAQlROv3IILL/qt9TC2zhRweFZHMC0HlC//Qwpv+YU1cK0nlPBwhMdwNDPUb8gria/jXTVwgKkVcGryR3A6UNTv6pHJL/wzdXCo6ZXwQWVHsDiilS/DXghv2kl1sIX/lnBK1cfwOiQVb8ntx6/44bWwnMsXMH2BSDAf4VWvyk/HL9b29bCk+ddwQ2aIMD1TVe/OyUav2o018JQ+1/ByjYhwPPIWL9TIRi/kprXwlbtYcHD2yHAKpxZv8LEFb8q5dfCLhZkwXk8IsCJDlq/Sl8Uv2o/2MLxuWXBmsAiwE7IWr8AgBK/NqDYwlK7Z8GvDCPAs0xbv6lzEb9++NjCQY5pwehlI8Dg+Fu/0D0Qv9FE2cIfmmvBgacjwHF2XL9qWQ+/i6PZwt+bbcF71yPA4y9dvz/ODr+aBNrC4GZvwfD3I8AgcV2/5l0Ov+NP2sItF3HBjyYkwAzTXb+tvQ2/7aXawnIhc8H2LyTAkh9evyuuDb+2BNvCq8l0wWo5JMBqSl6/cpQNv7JJ28IIaXbBOCUkwOT9Xb/Vzw2/O5jbwpIZeMEBLCTAvp1dv8GXDb+/59vCq+V5wWIMJMDcv1y/JNYNv0003MLp3XvBEf0jwFYZXL+g4g2/iobcwrZOfcHe7SPAd29bv3ntDb9I3NzCAgh/wVTUI8AYj1q/pREOv84W3cKqZYDBRLIjwDXvWL9RHg6/I2TdwvwngcGoqSPAqtRXvwPsDb+Ztd3C+eOBwfbGI8Dyxla/YyUNv1T13cKquoLBWbMjwHffVb95Lg2/2DPewndxg8FsqSPAKVRUv0rfDL8TgN7ColWEwfuyI8COh1O/rHsMvxbG3sK/BIXBnKgjwExtUr8QUAy/lf7ewpXPhcHYoiPAA+NRv1o9DL/MokRAHI++wgigR8Bp7Xi/HGG1vCB8QkCxir7CkSxHwFKaeL+2K+682GtAQBuEvsJNp0bASZ13v8ahF72SnD5ADH++wtNsRsCuh3a/y6glvQ7BPEAber7CG4dGwM6fdb/O6x69tNI6QJV5vsKV40bAnE91v3ctCL32KjlA7He+wuNSR8BNB3W/krrZvJyHN0ANdb7CrBJIwPRwdL/ueHe84Ko1QO91vsJTokjAV4B0v8Uj1rtQpDNAw3S+wgHeSEBrjXS/eu9CO17YMUD6cb7CvU9IQIsddb8eCTw8CNcvQAJyvsKI90dA1Cx2v6x5iTzuJC5AZHG+wumbR0AuHHe/ScC2PKSBLECccr7CwlJHQP2NeL9gU9s8iIMqQGVyvsLlFkdAK5l5v45X+Tx4rChAtnS+wpfhRkDkqnq/NSAKPcJlJ0BedL7CFMBGQMsQe7+uiBI9Qh8mQONwvsLmvUZAWlF7v44lEz2QkSRAC3G+wtXdRkAQf3u/jUgLPXhYI0BgcL7Cl+hGQNzEe7+TsAg9elkiQMRwvsIiOkdA3JZ7v0Da6Dwk/yBAl3O+wn2AR0CawHu/KgnGPDZeIECGc77CHcRHQN6Ze7+oc6Q8GE4fQIt4vsI4FUhAeZ57vyZ/eDxWyB1A+3e+wlhrSEAtcnu/oQwjPIRGHUCxer7CKsFIQIlVe79i8Zs7VJocQJB9vsI2DklAqzp7v8BA0DjCEhxAk4G+wgfKSMAmEXu/UU+Ku6ieGkAGhb7Ck5lIwFyIer9kCeq7DD0ZQPqIvsLkVEjA5oR6v7z2OLwYjxhAcoq+wh89SMDEznm//C9QvHB9F0CvkL7CMvhHwPyReb+uE4q8DA0WQMmRvsL78kfAsWR5vxuajLzcghVArpa+wu35R8AtbXm/BS+JvECyFEBEmr7CVvZHwDkheb/n34q8GHgTQBygvsLeI0jAEU95vxjoaLxkZhJAfqS+wvRRSMB71ni/JD87vJoJEUDPp77CWY1IwObyeL+WtQC8PIMPQImuvsJpzEjASAh5vwYOhbtUdQ5AkK++wnkGSUA9I3m/DBMUOhrNDUAHs77CBNVIQB8seb+DOGg7zEgMQIi6vsIwdUhAE4h4v41pGDzU1wpAmr2+wrkvSEBGDni/LqhcPMC9CUAdv77C+7VHQPUOeL+zQqo8hF0IQNXEvsIJd0dAorZ3v2gcyTzQxQdArM2+wqcyR0Bq93a/lmjqPGgSBkBWzr7CbtVGQAwtd7+1KQw93HgEQCjRvsIxnkZAOPB2v2yqGT3cFQNAwdS+wotZRkDgnHa/w2sqPdZAAkAq2b7CghBGQNmQdr8nVjw9LIAAQHXbvsL220VAZit2v18UST0oEv8/Rdy+wsGuRUCQGXa/dCRUPWBa/D+A2r7CjoNFQJOedb87h1490Cv6P5jevsKtcEVAmnN1v1EUYz3AZPY/392+wsVnRUAyNXW/lyZlPfCB8z/Q377CX0NFQJnMdL9J3209cIzwPybivsLdZkVAtb10v8coZT1Af+0/feC+wnhwRUCILHS/QY1iPWBw6T8j4L7Cj5RFQF8zdL8pvlk9aBPmP5zcvsIrpUVAaMVzv7p/VT3g/eI/dt++wrHpRUBaznO/t8dEPfiZ3z+8377C/Q1GQOa7c7/E4zs9sCHbP7vevsIgYkZAfQx0v2N0Jz2QvNY/PN++wj6ZRkCcT3S/kBIaPTB70j/T377CdtpGQJKQdL+KNAo9MGTOP0XXvsJXOkdAObh0v5eZ5TygCco/jta+wjGPR0BbRXW/d0+8PKCaxj8Y0r7CVsxHQGzndb8Xk548GCDFPwjRvsLiLkhAtBB2vxacXDy4Tr4/Fsu+wiBtSEDQ4Ha/0tQfPOBauz92x77CldFIQNGydr9Qj3Q7cOG1PxK/vsLfDUnAySZ3v0A0+bjgUrQ/PMC+wmnGSMCgBHe/rU+Qu7g4sD+jt77CdYZIwKgWd785AQe8KMasP76yvsItWEjAgAR3v551NLx48Kk//bK+wnQcSMDQXXe/30xvvAixpD9Yqb7Cm+BHwHE8d7+6CJW8SHqhP2ynvsJ6tUfANvZ2v3Yjqrzoepo/kJu+wtp/R8AUsXa/GGDEvCBblj8emL7CMltHwPdWdr8sOda8dEsPwSfJvcI36EJAXUR7v/VIwz20jRHB6L29wj6qQkCaXXq/HJ7KPXSrE8FAsb3C56NCQHNgeb+bAcs9QMQVwVynvcK9gEJAzrp4vxcXzz2+6BfBTZ69wuoyQkC2H3i/RnLYPcANGsHImb3Cd9dBQLwmeL+Kv+M9diccwWGUvcI6iEFA1R54v/GE7T0KOx7Beo69wj0BQUCGqne/Pvf9PZ5hIMFYjL3CMrhAQKeYd79YeQM+WpIiwSeIvcLWYEBAo4J3v0HYCD6MoyTBoIK9wrwbQEBzqXe/pCcNPqTIJsHEfr3CZQFAQMAqeL/u7A4+CNMowVR7vcLPwD9AaF14v3f6Ej483irBRXm9wt+3P0CB23i/Aq0TPjT6LMH1db3CDZg/QNU+eb81wxU+1AMvwYh0vcLHej9ACpJ5v+OsFz7G+zDBOHG9wuBWP0D5gnm/xuIZPtbdMsH+ar3Cp1A/QPSAeb8HRRo+TsI0wfVlvcLnTj9AKnV5vzxdGj4QmjbB7WK9woZCP0BoiHm/1icbPg5SOMFIXr3CW3o/QLBCeb/Umxc+rCQ6wZJdvcKwlj9A2kF5v7PZFT6YzjvBS1q9wuqxP0BK0Xi/gAgUPnSAPcGkW73C0tM/QIeReL9Z3BE+9Fw/wUJXvcLPC0BAwTN4v41KDj7y5EDBEFe9wmY4QEBXyne/aWsLPoqgQsHsVr3COmJAQB9id799uQg+tDFEwQZYvcL1jUBANNt2v4PiBT7ACEbBXFa9wmWmQEBqMXa/IDQEPuLcR8GoV73ClNFAQNgHdr/WgAE+lmtJwVRWvcKK2kBABCF1v2G4AD5wN0vBylm9wtECQUDeo3S/Gz/8PXQCTcE6Vr3Cn/FAQOtadL85Nv49YIxOwcBYvcLe8EBAMUx0v1tG/j0GQFDBxFq9wvfgQECN9nO/5wYAPqQJUsGXWr3CDLFAQAMIdL9A+wI+5slTwYJdvcKAg0BANYdzv6mjBT5MlVXBOF29wq8+QEDAoHO/4eEJPiRnV8HoXr3CwwlAQBxyc78LEw0+Li1ZwRZdvcINvD9Au6tzv//mET5g2VrB4Vu9wrOHP0A9XXO/fgUVPkCKXMGpXb3CTyE/QJOkcr8rExs+RlteweVdvcJt5z5AzDhyv2l8Hj4sPGDBf1m9wtR1PkBhFXK/+GQlPoTqYcGBVr3CGTI+QJ7wcb8ffik+IotjwQdYvcKu6z1AyCdxv/SHLT74g2XBp1O9wuqOPUCAZHG/XEozPpZSZ8GTTr3CWFI9QOY4cb+j7zY+FjJpwWRJvcIY+zxAGOpwv00pPD6G92rBmka9wsSwPECr4nC/PLNAPlzPbMHrPb3CRGM8QAZHcL9bNEU+VItuwWsyvcJqCzxAnC5wv7CKSj46i3DB0Ca9wpi+O0CBj2+/B/xOPlBEcsHAHr3C1n87QJ4cb7/WoVI+9DV0wWISvcKYUztAnOluv0VAVT7cMXbB7Ae9wg3kOkB5OG6/ScNbPrYXeMGM/LzCncc6QK8Cbr/ZZ10+UO95wf7svMICjzpAmPlsvydkYD5AS3zBxeC8wphxOkAz+2y//S9iPhYqfsGRzrzCfSs6QIpYbL90KmY+Mg6AwRzBvMJ5HzpApjtsv0fYZj6KFYHB6La8whX5OUDI4Gu/8QNpPlIsgsGPo7zC5ug5QEn6a79iDGo+2CqDwR+TvMK4yDlAUOtrv/36az66OITBpIe8wta3OUD842u/tf5sPgpMhcFwbbzC/8Q5QBfCa783IWw+tXmGwaldvMJHxTlADxhsvyZGbD6CfIfBAEe8wu+vOUCETWy/7KxtPqpJiMHeNbzCpL85QIopbL92pmw+qpyJwfAgvMJxwjlA/MJsv3jEbD4Mp4rBtwq8woLdOUDHVGy/KelqPu7ei8H+8bvCcuE5QDtmbL8CtGo+CbWMwUzku8LR9jlALP1rv8M0aT7v4Y3BQsm7wvwBOkAWuWu/eGZoPrrgjsHhs7vCQPs5QJZoa79KqWg+iOGPwQqju8IYJTpALoJrv8opZj5QGpHBAom7wkBCOkAZKGu/BjpkPj4gksGZebvC4Vk6QKajar9JjWI+wnCTwUxbu8JUgTpAOi5qv9ryXz7yd5TBUUe7whSgOkDrfGm/WsVdPvPpkEFIDrvCP+49wHX3c7+zWi6+r5uOQf4Nu8Izpz7AHkx0vw0YI75PYIxBzgq7wooqP8B3oXS/JSAbvt0ZikFPCrvCxcI/wN5Cdb8S9BG+FcyHQRgJu8LPc0DA4MZ1v9E0B77Ne4VBcgu7wsceQcBi5Xa/0u/5vXYrg0EQDbvCTNhBwGmEd7/3W+O9YNaAQT4Ou8IntELA9Kl3v6RQyL2uCn1BNhO7woZVQ8A7G3i/AJa0veUmeEHnFbvCMglEwPVNeL8ef569cpJzQVEWu8LusETAFdR4v8f1ib0gtW5BABi7woYjRcAiyXm/Ah14vVACakGYGrvCI8RFwLRger/aqFC9qCZlQYgdu8ImH0bA72h7v5SFOr3CWGBB/R+7wjuJRsDjY3y/jYkgvXJzW0GzIrvCSv1GwMxWfb+z9gO9t59WQUAju8KkZkfAMxl+vxHU07zi7VFBPiG7wlyxR8DIv36/qNOuvLI3TUFYH7vCH/tHwOJff792M4q8EoBIQbEgu8JqRUjArROAv/OASrxG4kNBjx67wuJQSMC1MIC/0R0/vLxDP0HDILvCf3BIwPNogL9UnR+8Q5Y6QVEfu8Lfi0jAF2qAv7EyBLzw9jVBziK7wmufSMBFb4C/VEDhu5oMMUGbILvC2KVIwAFbgL9OUdS7zpUsQTMhu8KLtEjANFGAv97YtrvezidBkiK7wue+SMD0PoC/vQ+iuy83I0ERJLvC9rtIwKELgL/R0Ke73WEeQasfu8Ia00jA9K9/vwndcrtzghlBCCC7wjTaSMAAp3+/8nVWu/PVFEHUHbvCMAFJwAgRf7/ANmq65QQQQZofu8I/C0nAUIJ+v2j5krnXMwtBnRu7wmrRSECVan6/aP54O8CIBkE4G7vCvbVIQF1Zfr9KprM7jMMBQcUbu8LMa0hAuwd+v0JtIzxQtvlAXhW7wpwhSEDl+X2/CE9tPEP970CSFbvC0M1HQMVyfb8+OaA8FmLmQCYRu8LPXkdAckZ9v/1h1zwuydxA4A27wmQDR0A4pXy/m0QCPajo0kDPB7vCv5JGQD2XfL+gPB498xbJQC4Au8KQLkZA4517v6DGNj1Bzr9AsP26wiixRUAXhXq/P2pVPb/5tUBJ97rCQ1NFQCb0eb8DZGw9XqurQB7uusKQyURABUh5v60Khz1WaqJA6+S6wtFbREBo63i/fH2UPcbgmEDJ3rrCMfJDQIjqd7+GNaE9dL2OQO/TusJ9ckNAdP53v4H5sD2144RAzse6wmASQ0AqqXe/TLO8PUpvdUD0u7rCA5xCQMZad79zKss9AJdhQEexusLFO0JAcVl3v9IF1z0mOE5AL6K6wgPUQUCit3a/l4bjPQSUOkBqkLrCwWBBQLygdr8/rfE9Ao0lQHZ/usIWBUFAaQV2v6ip/D0g5BFApmy6wuqqQEAkrnW/bsoDPiCM+z+UWLrCL3BAQDKRdb9yXwc+IFXSP61HusJl9z9AM7h0vwOQDj5wU6s/fDe6whnMP0BRyHS/pT0RPsj8gj+6H7rCFI8/QGi4c7/TrBQ+sBwqP8wNusI8cD9ArWpzv715Fj5AI7E+aPW5wjIxP0Ci8nK/ODIaPgASKT2037nCLCw/QCzCcr8Dcho+oC+WvlvSucI7Cz9AZjhyv1tLHD7gWiC/97i5whcEP0DjWHK/ccUcPvC/br/anbnCHPI+QNYjcr9nzh0+CFegv/qOucLR5z5AJvtxv61iHj6Aw8u/x3C5wo8GP0CDyHG/YHEcPnjB9797WLnClBU/QMEAcr/Wlxs+4OYPwJA/ucIXFj9AGAxyv2uTGz5Y6yLAbiq5wrMpP0DS6XG/t1UaPnzoOcCBEbnCZDc/QNJbcr+Mohk++PxOwE75uMKHXz9AIydyv9IdFz4IAWXAYNu4wsdmP0Ct8nG/c5wWPhC9d8ACzLjC8Yg/QE6Tcb9faRQ+7KCGwG2tuMIllj9A6Flxv/iJEz7g6ZDATJS4wqaGP0BQInG/fmsUPgCpmsCOg7jCpr4/QBIgcb8iABE+2tmlwNxjuMJy0D9AgfFwv7ncDz5MFbDAF0+4wozaP0C8dHC/ExsPPtzWu8AZLrjCT/A/QCz2b7+Wow0+MAHGwJAVuMJZ+D9AnXdvv1sCDT4rFJjBBf26wjh0PUAVV3u/CIs4Psz1msEPzbrCLeA8QOYcer+EUkE+raqdweOZusIUpjxA4WB5v8ioRD4uX6DBcWm6wkNRPEB4Z3m/9fNJPpgpo8GMOLrC1OI7QNYyeb9nwFA+QualwUcMusLQejtAvI95vyVkVz57rajB5N65woszO0AVhXm/ltJbPm91q8EasbnCobk6QETheL9fJmM+1D2uwe6GucJDgzpAbTB4v2I6Zj4MF7HBclq5wuEpOkDZYHe/bWxrPvrRs8EUMbnCy885QAeCdr+mnnA+EqK2wQ0DucJAlDlAiO51v5ULdD5qabnBUti4wiMqOUC8GXW//Tp6Pi4ovME2rbjCYgY5QCNVdL+rEHw++Pi+wSWCuMKsvjhAAwt0vwEvgD7UxcHBZla4wpt9OEBoynO/ByOCPnqRxMFsKrjCcyc4QE2Kc79zvoQ+REDHwYj9t8KF+zdA1Jtzvycghj6298nBF8y3wlLEN0Bat3O/at6HPmymzMHcobfCho43QEzxc79Amok+qErPwdlvt8LClzdArAl0v3dXiT407NHB9US3wm2NN0B0O3S/bbeJPu6e1MFaGLfCYoE3QML/c7/dBoo+qjzXwXTutsLhdDdAr8tzv/pbij5IDdrBdb+2wm2iN0Cnk3O/N+KIPrCj3MHklLbCt643QGVTc7/ubog+ZGLfwadptsIbwzdApRlzvy69hz78BuLBDEO2wi/vN0CWv3K/5kaGPpLV5ME0EbbCoAo4QD5+cr+0W4U+8KLnwQLmtcJTLThAZNJyv/hehD64VerBZrm1wrhIOEBoinK/wnKDPvgw7cEskLXCNWg4QDJocr9HcII+av3vwepetcJWYjhAy51yv9esgj6+vvLBVjW1wgONOEDB8XK/e3CBPkaO9cH9DLXCeXQ4QG3ncr9QMII+nnz4wdrXtMK1YThAfwtzv5TOgj52ZfvBhq+0wkpSOECZrHK/5S+DPuJG/sEsf7TCiik4QAbbcr9Jf4Q+ro0AwmpPtMJkDDhAEFhyv4BDhT5CDALC2By0wpfTN0CJfXK/8w+HPrKHA8J76bPCw6M3QCemcb9nUIg+B+sEwia3s8KyRDdAbZBwvxH0ij6lYQbCR4ezwhIhN0CYKHC/o/CLPur2B8JJT7PC/M42QCOyb7/fV44+ymEJwjcVs8KlhTZAhHxvvy2MkD6M3grCdt+ywr05NkCojm6/B56SPvBqDMIYpbLCEdg1QDCtbr/0qpU+CugNwjJpssJDjTVALk9uv7Ddlz53dQ/C4ymywnkhNUBo922/uxabPqf+EMIm77HCRL80QOXTbb9iFJ4+4XwSwimrscK2PTRAiOpsv7nJoT7SBRTC712xwueiM0BKpmy/C3ymPhagFcJgGrHCQxszQOm1a79vW6o+kCQXwvfQsMIhljJADg9rv4o+rj7FsBjC9oWwwgwaMkDQnmq/huyxPhVPGsL2OLDC8UYxQB93ab/aBrg+9dMbwlTsr8JwtTBAcdRov0pJvD7UWB3CjZuvws4bMED4I2e/kmDAPl8aH8LmSa/C4o0vQBHHZr9fnsQ+sp4gwn7yrsLiyi5ANHVlvzgbyj4MKiLC/ZquwoQ9LkCK32S/YDvOPt/DI8LsT67CDqYtQEq6Y79ebNI+3WMlwmP0rcI0CS1AMzpjv0oO1z7a4ibCTpOtwv1gLEDwRGK/AtXbPkp4KMLCRK3CmdorQM5wYb/InN8+4B4qwgLhrMKcZitAKmBgv3O04j6SzivCPIeswsfrKkBa6V+/m0nmPlhVLcLWI6zCN10qQLkTX7/aTOo+l78uwgLHq8L65ilAKCtevy+E7T78iDDCpGirwhKPKUBAC16/6S3wPu4mMsLKBKvC1jwpQKM3Xb/iT/I+6tszwmmbqsKS2yhARmRcv9Xn9D7JOzXClESqwlebKEAuZ1u/nmL2Pg3wNsLX3KnC91IoQFXJWr8dTvg+B3Y4whB8qcJX+idAEgNavy2k+j6a+TnC7xqpwuX1J0C5vVm/uqL6Pia6O8Jns6jCONgnQBVJWb/fT/s+1Uk9wjVYqMIRyCdAyVxYvw5S+z5S+z7CK+qnwo/BJ0Di61e/WUn7Pjh3QMKaiKfC8rsnQG3dVr+V5fo+AQSuwmwCJcI32RdA3FA+vz2WND88ya7C/skjwiztF0CizEC/EkY1P3p4r8KEjyLC+zgYQBFbQb9WSjQ/LCawwjtbIcLwNxhAChNBv68xND/f2bDCViIgwnILGEAWkD+/Iks0P0GFscIN9R7CFd8XQCPtPb8aVjQ/8zCywurJHcJrsBdAOgI8vwdMND9D3LLCW5scwiU6F0BRpTm/djI1P12Fs8KwbhvCBAUXQOtfN78mGTU/SCm0wkBBGsKqnBZATWI1v8roNT8Sy7TCpBYZwvI8FkAqwzO/1ro2P/JstcKu4BfC5PEVQK6VMr9UaDc/Iwi2wkK1FsLVgxVAHggxv9N3OD+6obbCsoYVwk9pFUCeDTC/7HY4P9lGt8KYVxTCcz4VQEN4L78Q4jg/ktq3wsMoE8JKGxVAdL0uvyEeOT+scLjCKv4Rwvb3FEBEJC6//2g5Pwz/uMKW2hDCS+EUQOwpLb8/Vzk/g4y5wgymD8L4xBRAYvcrv0RDOT8FF7rCHIcOwryDFEA5yCq/hcI5P5icusI+XA3CU2gUQP1/Kb8RoDk/vBa7wgY/DMJ7PBRAY9Qnv26SOT8korvCKCgLwp4qFEAsMCa/riA5P+savMLgEQrC7QEUQBFSJL9v7jg/Ip28wgfzCMKJMxRAwLYiv+R3Nz8UCL3CzuMHwrkkFECMwCC/WNU2P86GvcKE2QbCCC0UQElrH794HzY/XQK+wnXYBcImNxRAdBoev8lkNT+qb77CsMMEwrZWFEA8KB2/UYE0PwTgvsIBsQPCA3IUQE4hHb9hFTQ/T0y/wpOxAsJ0ehRAXF0cv5+fMz9rwb/COrABwm6PFEAHuhy/iHczPxQrwMJMpQDCKKEUQGD7HL/1TzM/M5nAwlZO/8HfyRRAYJcdv2r3Mj+IAcHC+Vz9wR+xFEB3GB6/XY4zP4d0wcJxQfvBMqgUQBihHr/36zM/Pt/BwqVf+cHdmhRAQ4kev/EUND+qR8LCCVr3wamMFEBtSh+/pJ80P7GgwsJRVPXBuVAUQNg+H78CgjU/RwnDwtBU88EBERRAlmcfv1aKNj/hacPCuGPxwevUE0AzPR+/dGA3Pwu8w8L0Ye/BWWYTQOisHr9wzTg/5RXEwpKT7cFBMBNAgFkev096OT+KfsTC8JnrwXbqEkAhKB6/inM6PwnbxML7nenBKo8SQLLSHb+BsDs/9S/Fwu+/58FzRxJAvQ8dv2VvPD/FicXCNtvlwc32EUCCdRy/WWM9P87ZxcLc/OPBipIRQHDlG79qqD4/oS/Gwhgf4sEaXhFAvH8bvy5GPz9heMbCNy/gwZUJEUD/uxq/aDVAP/fOxsLPR97BvLgQQLgcGr8AJ0E/XBbHwlA13MEAFhBA0iUZv8QvQz+gaMfCkUjawTnGD0AjcRi/bhJEP2C2x8L4StjBfGIPQOCeF7/jNEU/RfvHwutW1sF65g5AvqgWv+SkRj80V8jCYDnUwR9NDkDFUhW/tlhIPyqXyMKhU9LBKbgNQGMDFL9Q/Uk/PNfIwttc0MFGEw1AGugRv3h4Sz+TL8nCSDrOwY9yDEAZjBC/rUFNP7JqycKqH8zBd60LQORPDr88JU8/3a3JwqYcysG+7QpAFowMvzEvUT+a6snC0xzIwYomCkCRDQq/zvBSP/g5ysI3FcbBSm8JQBIrCL8Gw1Q/in7Kwpvsw8HxpghAABcGv0m7Vj8cxcrCQu/Bwcf4B0AUsgO/ShtYP6AKy8J/2r/B7RAHQJi9AL+rClo/qkLLwrm0vcES5QRA2hj9vjimWz9piMvCwKy7wQ6FBEAVpPm+CBldP+20y8Lej7nBl7ADQPQt9L7MZl4/jRnMwhaDt8F3agNAoeTxvsI3Xz/LUszC6j21we4fA0CS2e6+6JNgP2CczMKXBrPB7nMCQKSl6r4yXGE/GL3MwpL8sMFO0wFA1MTmvoQFYj9xDc3CtdyuwYq1AUAaYuW+A61iP0E7zcLLmqzBSIYBQC+4476LR2M/0HjNwryAqsFVRgFASYTivtlAYz8Pvc3CEmKowRr+AEBaK+G+YDZjP3gEzsI4b6bBLI8AQE7u3764fWI/XzXOwgNZpMFOfwBAQ+jfvoJEYj9UU87CvBOiwZXW/z8OsN2+3MBhPwYrDULlw7HCOIwzwD2VaL9dzqW+4oILQgfZscKzvDTAXLxqvzMrnb7t6AlCaeqxwhSvNcCPxmy/lFqWvgRGCEI2/rHC5pc2wCa6br9hw4++W5sGQrsQssJegzfAlD9wvxnuiL7+7QRCDSeywqZTOMCyG3K/BP+CvqdAA0KnPLLCBUQ5wIkUc7/Snne+W40BQg5SssL6NzrA1H1zvwC8aL7iv/9BmGuywrD4OsBL83O/KQldvgoz/EEogrLCpcE7wLMIdL8yqlC+e9f4QaWYssKEkTzALXp0vywFRL6aUPVBFq2ywogoPcCgenW/bxU7vqrf8UHowrLCae49wFwbdr88Fy++W0vuQQjYssIZcz7A8Vt3v61QJ75wyupBUeyywngOP8CZp3i/sRwevigy50FUALPC07k/wDD/eb8/4xO+OKHjQdIQs8I9VkDAo0h7v5qHCr5nIOBBnx+zwubYQMB8Y3y/4q8CvqKc3EGwLLPC6WdBwLqCfb/IF/S9VA7ZQXw9s8Jb+UHA4PF+v56T4r3vkNVBEkqzwoZXQsCepX+/LBfXva8W0kHtWbPCK8VCwKM6gL+krcm9BoHOQcVks8KcMkPAKWyAv8AYvL3M8cpBA3WzwlmXQ8BsmIC/z5OvvZRDx0Erf7PCzvZDwEilgL8BoaO9xMjDQQ2Ks8JcY0TAMMeAv1IVlr1GJ8BB+pWzwijGRMDDyYC/nq+JveSPvEF5obPC1w1FwMC3gL9Np4C9zuy4QRals8K8ZkXAbqWAv4Xvar0CObVBBq2zwsS0RcA8sYC/92ZXvR6ZsUF6s7PCrxhGwICEgL/XNT69WfKtQQO9s8J0WUbA6UqAvzfTLb3LPapBj8GzwsPHRsAWVIC/9jkSvcuipkGQx7PCBflGwJpSgL+H4wW97/CiQVvOs8J3ZkfArjKAvyff1Ly+L59BtsyzwlW2R8DtKYC/vOCsvMl9m0HM07PCPBZIwDnmf79kk3m8ws+XQXPUs8JIeEjA/Jl/v3Z0F7y5I5RBvtazwoTQSMDEyn6/TcF8uwRdkEFJ1bPC+/pIQG2Zfr+GiqY6ZZqMQS7Ss8J5pEhAkz59vyKb1Tt6AYlBANazwslOSEARB3y/PZI/PKBNhUH707PCGQ9IQG5Oe79OaH48uWaBQZbRs8J0t0dAS0x6vzNKqjycqXtBH86zwp1pR0DQ2Xm/6JjQPL1adEFozbPCvShHQGbKeL9+JfA8o8dsQVbJs8IB2kZAhr54v6F4Cz3wRmVBacOzwjuhRkDkR3i//FMZPTOLXUFuvrPCKlNGQBrsd79Tbyw9M/hVQRW5s8LdIUZAYQx4vzWeOD0Np05Bm7CzwpPiRUBqX3e/T+9HPe45R0GSprPCd5tFQHhUd7/mZ1k91mM/QYycs8KraUVAD992v/NxZT3A3DdBTo+zwnMtRUAsrna/pSZ0PaZSMEHIg7PCMhdFQCy+dr9Jpnk9Hq0oQWB7s8K7v0RAB/11vzZehz3mWCFB3HOzwg2xREAhPna/Ij2JPTXBGUEHY7PCi41EQAhcdb86WY09YXgRQVxZs8Igi0RAFQt1v0aOjT01wwlBCkuzwhpoRECs6XS/q86RPYyGAkFePrPCGntEQFDIdL+RcY890in1QIk6s8KVaURAWGB0v194kT28m+VAuy2zwnF5RED7wXS/y6OPPc7p1kBkGbPCqG5EQDiodL9S7pA9rN/HQNsVs8KAdERAKp10vwg0kD204LdAOASzwtibREBbkXS/J2CLPRcOqEB387LCVbFEQHEGdb8g34g98DCZQJTpssLdvURAQkJ1v2Fmhz1OZ4pAQuCywhLQREDwc3W/fziFPdKVdEAK07LCBtpEQP7xdb93IYQ9tDlVQHzJssKq/URAu0R2v4+ufz2azTVASLeywhUIRUCpJna/rhB9PdbUGEC4tLLC+iJFQEYxdr8cfHY9sIXzP4qissIHJ0VAUEd2vzCIdT04+rU/h5iywm0LRUCwgHa/zGt8PfCUeT+WlbLCakFFQLKodr99PG89ABnyPrCBssJXQkVAFc52v+ETbz0AOBW8H3qywjI6RUBQv3a/UQ1xPQBYCr8PZbLCDDVFQOSSdr/BO3I9kOyAv+RZssLHKEVAPZN2v2g/dT1UySfCm2quwg1FNEBR83K/AIqjPhXlKcKG+K3CeWYzQGY3cr+mQao+KtsrwuSArcIH+TJAc1Rxv+BfrT4Gzi3CbAutwiZnMkDonHC/La+xPsHOL8I6lKzCpMYxQKglb7+XK7Y+SL4xwk0jrMLoJzFA1Ottv76ruj7HtjPCAbKrwlymMEDPDmy/z/+9PpKqNcJ5P6vCFd8vQGKzab9ORsM+uJo3wqLQqsLfZS9AaS1nv80Exj4QijnCWV+qwomxLkBxsmS/d5LKPrxwO8Km+6nCkPctQE0yYr8KP88+EFg9wlSGqcIIUC1A8g9gv1540z7IOz/C1BWpwsuULEAO6l2/BkLYPm4OQcKVo6jCcCAsQF3ZW7/r4No+L/BCwlYxqMLyjytA95Fav5Wz3j63w0TC+runwvIWK0AmaFm/99jhPqCaRsKRSKfCeoYqQIWjWL/04+U+WFlIwt/UpsJnJCpA+i9Yv7Gq6D4pFUrCB1emwknBKUA4+Fe/GZfrPhrOS8Ku4qXCgEIpQO7MV79TY+8+lX1NwuJkpcLmAilAEspXvyhV8T5XFk/CCu6kwvKzKEBEnFe/9KnzPurUUMIadqTCumQoQDonV7/72/U++W5Swnj+o8JlBShABnFWv/Rp+D4+H1TCa4Kjwg/8J0ClzVW//lz4PgewVcL7CqPC6rwnQGn9VL8y3vk+W1lXwpaQosK9fydAN0hUv0pd+z5K7ljCEB2iwltgJ0CzeVO/OuT7Pj6JWsJ6maHCMzknQErCUr/Os/w+xCdcwlgbocKHCydA7sxSvzsf/j5OtV3Ca56gwoziJkAkBFK/ifL+PjhWX8K7IqDCRKgmQEqfUb+qQQA/cuZgwj2fn8KwaSZAb1JRv2IhAT/4fWLCySKfwjpqJkDMOlG/tRgBP7QFZMJAqJ7CiQgmQPvBUL+HdQI/hKxlwr8bnsKkwSVACUVQvwVoAz+sRWfCr6Cdwh2CJUCLPU+/6hUEPwDTaMK6Gp3CfTclQKfvTr+5IwU/jE1qwtaTnMLn2iRAsq5Nv0oxBj8q5mvC6gmcwlFzJEApA02/ZpQHP3R2bcLSf5vCvw8kQCVsS79YoQg/ZuNuwl32msIQeiNA/51JvylfCj84YHDCMnGawoMrI0AiZ0i/FDMLPxgHcsLJ5JnCycUiQA0hR79eWww/pHxzwihTmcL3VSJAszJGv9bEDT+wAXXCSsWYwu/bIUBarkS/yyUPP5SKdsLQNJjCWmwhQEThQ7+5lhA/NgV4wlall8Ix/CBAiLVCv6rqET8Ej3nCFBCXwumEIEDZ30G/NXUTPygHe8JhgJbC+BcgQD3+QL/q0hQ/BHp8wljolcLieB9Af5E/v2TEFj8y+33CkECVwqGwHkCKmz6/Jn0ZP4Z3f8LqqZTC/CMeQDh1Pb8VPBs/7nmAwtoKlMIcjR1A7Y08vxQ4HT9UMIHCvmuTwmnzHEDqkTu/MjcfPwb6gcLgxZLCCxIcQAQ0Or+tKiI/zK2CwpgnksKkWRtA0xg5vz+VJD+FYYPC1oKRwnOkGkB8Dze/PZgmP+IthMIj1pDCBvIZQGhdNr8nESk/9dmEwggskMLPCRlAioY0v8vrKz9Rj4XCHn6PwvhZGEBNhjO/czsuP1lEhsJm347Cd6QXQGbeMb8oXTA/fASHwjwzjsLl7RZAr/Ywvy/QMj9rsIfCqXmNwugeFkD6hi+/U2s1P3tiiMLu24zCiooVQN09Lr/sLDc/9yWJwr4kjML83BRAU1ksv0MOOT9g3InC53eLwpU+FECjaSu/sRs7PxaOisJ+xIrCn6wTQHlTKr9c5jw/5CyLwpwTisKJBhNAjfkov0XiPj8bAozC6GOJwhG6EkBkxCi/q/s/P4O0jMLqrIjCeFISQBsZKL82TEE/7X+Nwojwh8Jz3xFAFRwnv86kQj+QC47CqkyHwiugEUCeMCa/mzVDP8vTjsKzlobCOmMRQAguJr+XKEQ/znuPwgLkhcIYFBFAcN0lv5xART/wJpDCwS+FwvAPEUC82yW/hFBFP1/1kMIge4TCgAgRQBsmJr80kUU/n6uRwhnTg8JSFBFAZrElvx4rRT/pZpLC3RSDwosfEUAkLia/WDhFP2AEk8IbYoLCRCYRQBWIJb/Nz0Q/PROTQiWvgMJeJt++3IZmPyR4177HZZNCttKAwkQC377WHWc/+ZbXvtuyk0Ls94DCiZPdvvjHZj9YA9a+J/6TQlIfgcJjQtq+5RVmP4Zo0r5ySJRCGEWBwvWR1r4rqWU/YZDOvpeSlEKOaoHCfCTVvn48ZT87+My+n9+UQsCPgcJ0ZNO+ZYFlP7JZy74pLJVCA7aBwgFh07745GU//H/Lvnh4lUKO3IHCHGTUvuZZZj/uscy+YsWVQhMDgsJO8da+tqlmPzpcz74cF5ZCdC2Cwo/h2b7wxmY/zFTSvrdhlkL7U4LCq4TcvliPZj+p3NS+tLCWQlh7gsJ1XN++eIpmP6mv177p/5ZCZqWCwg/w4r5tFWY/cwzbvjhPl0ILzILCb+vlvjvnZT9h8d2+F52XQof2gsK38Oi+u3tlP0fE4L5w7JdCVR2DwjAg7L7XQmU/1NjjvgE8mEKdRYPCnazvvrIOZT8lTee+g4qYQpJrg8LJjvG+AidlP9A76b4W15hCppODwqcy9L6FF2U/ItrrvlIimUIcuYPCzCX1vntbZT9Y7+y+22qZQjfhg8JlSPa+DpVlP40v7r67uplCOgmEwunB9r7OuWU/zrvuvrUEmkI1MoTCmfT2vq4EZj+hE+++Gk+aQixVhMJhnvW+TCNmPzbK7b7omJpCQ3yEwjJu9L5TImY/u5fsvtHkmkIMpYTCCpHyvlFBZj8Ix+q+tzWbQqTLhMKoU/G+JNdlPyRV6b6VfZtCC/KEwn+o777XgWU/7X/nvsjFm0LrFYXCa8btviI1ZT++eOW+aBacQq1BhcKxgeu+yg1lP4Mh476PXZxCameFws6E6b40vmQ/RQDhvumknEIpjIXCClTnvisxZD81kN6+Z/GcQnWzhcKD4OS+ueZjP+r9274+PJ1CL9qFwuHC475giWM/27favuyFnUJjBYbC/aXhvoxLYz91g9i+fcydQqsohsJzTN++rGVjPwk71r4lFZ5C7lGGwqrM3b7yQGM/3K7UvoxknkLFeIbCjnPcvssuYz+DUdO+W6ieQnSghsLiStu+bxljP+4i0r6N9Z5CNseGwrU12r4v3mI/dPfQvv49n0IW9IbCSajavs3+Yj+6dtG+ZISfQgQZh8KgJdq+qx9jP+gD0b50z59CNUmHwmh72r40zmI/WDXRvhoWoEK7cYfCe4TbvomeYj9WJtK+WlygQgahh8LLkNy+zEZiP+gI076QqKBCyMuHwhHv3b4yD2I/QUrUvlrxoEKD9IfCddfevvCVYT8W+tS+I0GhQrQjiMI0CeG+4WphP1MR174agaFCEEyIwgEW475fc2A/BKjYvvHFoUJUeYjC3ZXlvvgiYD+n+tq+8xOiQm6miMIDKOm+NIZfP5A43r6WW6JCSM+IwvjG677j/V4/MI/gvmqhokIp+YjCEZbuvqggXj9a7OK+4+aiQrchicJgBvO+8nhdP0T95r7/KKNCmUuJwl94977Mqlw/Lfzqvmt5o0Iac4nCkpv7vqTmWz/bsO6+x72jQtyhicIbOwC/ejpbP+kl8760+6NCtcyJwr/FAr9R5Fo/bQH4vipEpEIA94nCCEsFv3BXWj+utvy+yIWkQgIfisJA0Ae/7k5aPw7aAL9uyqRCwEuKwq2zCr/yc1k/6n0Dv/kSpUICdorCTJ4Nv8YmWT/cUAa/jFalQuyiisIkfxC/Go1YP0sFCb85maVCjM2KwoeXE7+6Blg/rPcLv+nbpUK4+orCTggWv0a7Vz9xVA6/ThemQmwki8L2Dxm/d0RXP/E8Eb+gZqZCplWLwhQjHL9XpFY/UyUUvz+epkJeiIvCpKsev9KsVT9WZha/Y+KmQi2ui8L+PyG/aP5UP3vKGL8RGadCQ+GLwg6cJL9+JVQ/jesbv2hcp0JZCYzCzdEmvyFrUz8W7B2/qpinQq47jMJJRSm/8FVSPzsMIL9A2KdCoWuMwnteK78YdVE/uOEhv4YRqELTnozCiOYtv+Q/UD8DCiS/GU+oQtTRjMLa6C+/L+ZOP6SbJb+5j6hCAwCNwgf+Mb/AlE0/6EEnvzrGqEKUMY3CM4Mzv5NCTD+TUyi/PfaoQoBjjcIJNTW/GN5KP46KKb9JMalCmpGNwseHNr+eVEk/YlEqv3TdvMLf9+/B0sk0v8fsOT/f3iK/zj+8wsgV8sEw/zS/Zxk7PyKCI79EobvClin0weESN79fUzw/BgMmv6wJu8KCOPbBxx85v2zTPT8pnCi/WHe6wuxV+MGqOjq/MeY/P/9/Kr9A37nCFoT6wTqTOr/BmkE/Pn8rv/RGucIgvPzBnX47v0enQj8o0yy/0ra4wjz//sG9YDy/UB9DP+LlLb/yHbjCIKgAwqy/Pr+ZZUI/owQwv8KIt8LS0gHCe05Bv5y8QD9n9TG/4e+2ws4AA8K+1UO/mWg+P2yXM78CX7bCoy0EwqV3Rr/R3Ds/ZTk1v5nEtcJ4ZAXCyWZHvy73OT8jZDW/VC+1whegBsJtOEu/s/82P+cCOL9YnLTCyMsHwo/LTL/+MTU/nNY4v6MCtMIFEwnCSt1Pv8FuMj94vzq/WGyzwkpICsJ+PVK/+wMvP0OnO7+ezbLCmoMLwr3qVL9Ctis/j+E8v0g4ssKJtQzC+pRWv4xcKD+SDT2/PqWxwtv/DcJgp1i/AkIlP4+3Pb8MD7HCYiMPwlXSWb/gBCM/Cdw9v7B3sMJzaBDCsFJbv1EnIT+afT6/VduvwiuhEcKSnlu/JJAgPzODPr9JRK/CsNgSwh+ZW78L0yA/b5w+v5irrsKOBBTC0g5bv/0GIT8NLD6/gg2uwggxFcKxwVm/HXYiP/WKPb/weq3CLFIWwpx/WL8xdyM/v8E8vwLcrMJlcxfCalFXv5dzJT+2ezy/uT+swryRGMKuPFa/ce8mP4gTPL/uq6vCQagZwnDGU7/S7ig/w4U6vzULq8IAvhrCcvtQv5KXKz8a6ji/nGqqwj3DG8Lbok2/NDguPwG5Nr9m0anCoM0cwssfSr/LazA/qio0v24uqcKzwR3CMPlGv23uMj9TFDK/Wp+owjHAHsIcx0O/fhU1P5LIL79o+6fCgMMfwvjDP79Ppjc/RdMsvzpkp8IprSDCgqg7vy4nOj8iuim/4M2mwm6ZIcKxxje/4Bw8P4CgJr92L6bCBIoiwgFBNL++2j4/ZCgkv4OUpcKiaSPCdbEvv1ODQD9SPyC/XPekwlU+JMKbPyu/z+hCP/mzHL8sUqTCxSIlwu7MJ7+pI0U/1hAav8W/o8Ix8CXCbjEkvyYFRz8eIxe/8xijwvy1JsLmXiC/9RFJP+QJFL9vg6LCeHgnwlfoHL9GYko/vQsRv3XdocKYNSjCOu0YvwhdTD93vg2/aj+hwvnuKMIZExa/BIFNPxVJC7/5pqDCWJ4pwlnJEr/A504/lngIv68EoMKrXirCzugQv521Tz++3Aa/GHGfwkgCK8LLvA6/1nRQP/vxBL9W2Z7CFaQrwgktDL+NllE/qMACv+c3nsL/WSzCL2EKv4ApUj+qJgG/DqedwpT/LMK3bQm/wl1SP3lGAL/XGJ3Cfp4twk1oCL/BBVM/qej+vgmDnMI4Ki7CtiUIv4Y0Uj8t9P2+w+ubwgvMLsLCDwe/mI5SP2oF/L7SU5vC4GAvwtJcB79cKVM/7O/8vt+8msLt/i/CbpEHv+zjUj9SMf2++CyawmCWMMJKMQi/fb9SP41W/r6SmpnChjIxwo46Cb8Wd1I/TxsAv0wImcJIuzHC28UJv/jbUj+nvwC/xoKYws1YMsICRQu/rBRSP1AAAr8JAJjCR+MywrJFDL8D7VE/JvECv8x5l8JkdjPCdS4Ov8pdUT+8qAS/BuuWwogQNMIzNhC/5y5RP4qaBr+GaJbCf6w0wlLcEb+Qr1A/RRUIvzDdlcK7SDXCSJ8Uvz8WUD89oQq/fFyVwujjNcIQmxa/80BPP4pWDL/j0pTCDHY2wpgnGL8rEE8/qs8Nvx5JlMJ+GjfCRVAavw7+TT/lnQ+/B8qTwo6/N8JbrRy/js9MP/CVEb8eS5PCtls4wrhaHr+khks/SdYSv0u8ksIH/jjCcv8fv7IdSz/5VBS/HDSSwhmgOcLPFiG/jZhKP30+Fb9OtJHCa0g6wvIlIr/Wp0k/XPwVvypSkcJt5zrC+20jv0DHSD8c9xa/driQwmOeO8IfeiS/7OVHP6i1F78nSJDCRUk8whgrJb8uBEc/TxkYvwy1j8Lq3zzCcGAlv4qARj/1IRi/N0SPwu6oPcJ68yW/3hRGP/yOGL8YsIHCag2Xwr+2xr3a8HM/ly/CvW/wgML8HpfCY5DdvQbcdD9m9Ni9tzCAwkssl8JF/fO9ojh1P74i770s4n7CgTiXwr2PAL5p6XQ/ruL7vXxkfcL1RJfC3eUDvm7cdD8CNQG+utp7wqVSl8IqmAW+C/10PyzoAr5EUXrCg16XwiOGB74HXXU/DucEvirJeMLLaZfCYcYIvkwVdj9VUwa+QEN3wpp6l8JRiA6+CP12PxVADL6Gt3XCi4aXwn8YFL4YXXc/GdoRvrordMJonJfCOToZvg+qdz/BAxe+CqVywn+ml8I1lB++au93P0tiHb6WHnHC17SXwlx9JL5mVng/lWAivnCSb8JJx5fCyTQvvhg4eD879yy+zg5uwrzSl8Kh5zS+Ic94P1PUMr6YkGzCPeWXwpMpPL5L3ng/uBA6vlACa8LJ9JfCOnpBvlyeeD/oQT++mGppwskAmMJJi0e+CFV4PxYuRb4O6WfCQQ6YwqrPSb6kAXg/Uk5Hvu5lZsJqH5jCK/pNvhiDdz+RP0u+xORkwpUmmMJjMlC+bjJ3PzRTTb5iZGPC+TWYwitDUb4nlnY/ASJOvoLVYcLNQpjCWk5QvlSqdj9pN02+HlBgwm9TmMLchE6+l4Z2PyhjS77C2F7C3FuYwn1tS75yFnY/5CVIvvRKXcKNZ5jCAbNEvp4bdj+GfkG+wuZbwqNvmMLgcz6+dih2P/xUO74+VFrCqnmYwj/tN75/hXY/NwI1vn7dWMLPgZjCm/QyvvFJdj9pAjC+tnVXwkyJmMLOSim+OF52P558Jr7SBlbCQ4+YwiFHIb5EFXc/Jcwevg6gVMIflpjCrTIZvhM9dz+f2xa+5UdTwq6bmMJmNRC+xX53P0kMDr775lHC852YwuZcCL4G0Xc/ymEGvlCRUMLmpJjC6HIBvozIdz8qFf+9yDBPwnStmMKcZfG9XSt4PwX77b0D+03Cg6uYwjiU4b0Eyng/iqDevY+3TMISsZjCXhLUvZwHeT8sX9G97HxLwm23mMKkpMq9y4B5PzE8yL1TN0rCTbeYwtwKub06y3k/hey2vZX7SMLftJjCmJCsvRJFej8LvKq9vMtHwli7mMKu7KS9ZTN7P4t3o73dq0bC6LmYwvBWnb2flXs/OA+cvVl2RcKLupjCHZqUveKefD+dr5O9REZEwn66mMJN6429hAh9Pz4njb19LUPCXrqYwk+DiL3KB34/xQmIvcL+QcJ4uZjCpvmDvejNfj9rtoO9xvVAwqeymMLD/3y9Q6x/P/jqfL1Fzj/CUrWYwrjPgL1yWYA/nQeBvQLDPsKxr5jCckd+vR+lgD9sAH+91bc9wlexmMJOuoC9fDGBP9Jegb1NoTzCLreYwg/hg73Xq4E/NcmEvSqkO8LHtJjCCCSLvWcVgj/EU4y9eKo6woOwmMLntJC94UCCP4sKkr2xvDnCE6+Ywncdmr3IQ4I/S42bvZTNOMIKsJjCClKdvZFWgj/g1Z69vc83wkmumMIg/Kq9ZJiCP5TRrL2N4jbCLrWYwkTmtb3qe4I/C8m3vUgDNsK8tpjC4JzEvYgcgj8HY8a9SiQ1wle+mMI2uNC9YcmBP7xb0r18WTTC+bmYwpAh3r1xc4E//pvfvbZ8M8J3w5jCemLtvXifgD+lOu69hKwywmfEmMIQp/q9JOh/P5br+r0Z5THC/cuYwhayA74qWX4/uXMDvlIdMcL71ZjCTgkMvtFwfD+tRwu+NGMwwvnamMLVrRK+c1Z6PzBOEb7zmS/Cx+eYwpU3HL7rb3g/dDUavnXhLsIY+pjC8CEjvr0Hdj+RSyC+fxcuwqIGmcISWCi+ktRzP722JL6OWS3CxAeZwswZLb6u+nA/d2wovgujLMJQF5nC9vg0vqFAbj/TJy++uvcrwvommcKiHji+2o1rP7M+Mb4TFSvCNS6Zwkn3Or7+eGk/5zwzvgBXKsJZQJnCKPU9vrc7Zz8CSjW+upUpwmRZmcKi40C+ZCVlP49QN76r6yjC5GSZwsk4Qr6022I/4bc3viYKKML8eZnC/EVEvgEjYT9CAjm+IlknwmSOmcIkXkO+2X1fPzGFN75cgibC/JyZwmyAQr7+Hl4/+C02vozMJcKesZnCKCZBvirtXD/9czS+HmbQwow3U8EWKBnA9fNCvzwfMb/iKtHCSNlXwaU4GcCG9EG/v3cwvxH10cILIlzByxkZwBRYQb9mtzC/WMXSwhpSYMEixhjAGAZBv13qMb/YndPCGYRkwbjVGMBus0C/Zooxv0hx1MIoz2jBix4ZwMseQb+gjTC/8ErVwkP8bMFGsxnAu+lBv+aCLr/uJNbCTApxweqWGsC88kK/glErv84C18IyHHXB/mMbwI1SRL/Umyi/Ad/XwlvweMGpeBzAsBZGv/npJL9VtNjCNz58wbUNHsA9eUi/8Gsfv9yN2cK3v3/BkXgfwG74Sr/snBq/XmXawuG0gcG8MCHAgEpNv5mGFL92PtvCQ2WDwV/GIsDMxk+/ugMPvyYY3MJSGoXB1YIkwKrVUb8gwgi/0ePcwiGxhsHeDCbAenpTv/gnA7+Ssd3CQy6IwavMJ8BO5VS/w1X5vmN43sJ2tYnB2DkpwG5SVb/3ZO6+1UTfwm4oi8HQpyrAXKhVv7Vt4754CuDCKICMwbr3K8DKXVa/bpbZvgLS4MLF043B4hEtwGBAV7/HdtG+kInhwjoCj8EX7S3AJsBXv3cSy77uTuLCjUOQwbXmLsDkwVi/P/vDvm0J48L4gpHBWKgvwIUyWr+Gvb6+f8XjwpyCksGWSzDAM/Bbv/eCur5OduTCcomTwYjhMMChAl6/9Mi2vqo05cJiiJTBSTkxwC3EX79Sy7S+kOTlwjJ2lcG1lTHA9uhhv97Jsr6nlebCDk+WwezeMcBSB2S/jlaxvlpC58IR/JbBTYgywPSIZr8KEK2+4PbnwvXDl8EEwzLAcENov3Pfq77flujCb36YwTEwM8D4PGq/uzOpvp9F6cJGHZnB1a0zwMnNbL/PMqa+kfbpwoa2mcH4NjTAMwNvv22vor6voOrCGkqawdnTNMCfunC/Ul6evqhA68IhmprB4b01wCXgcr8EwZe+Ou7rwrU4m8EaVzbAGzV0v1dgk77GlezCgJGbweZxN8CO2nW/qgWLvjQ57cJB0ZvBP284wFTzdr9taIO+FNLtwjUKnMGHsznA8yh4v+ksc76Mfe7CzzWcwWgjO8BHkHi/kmlcvtwO78LMjJzBwJQ8wGN+eL91ZEW+hKDvwtqhnMGbIz7A5KZ4vzetLL6mQvDClYecwSDkP8ApJ3i/Z7sQvhLQ8MIVaJzBp/dBwN5PeL+E2N+9Em7xwisynMEbGkTAmJ53v+cynL3wAPLCPwqcwe5IRsAWpne/K9wuvXaT8sLYx5vB7cBIwPKjdr+0+Jq72DLzwmaAm8FJ/UZAdO92v0VOAj28v/PCNAabwQBzREA9O3a/dNqQPQRM9MKZZprBk9NBQCRKdb/07uI9kMn0ws7vmcGTET9AijZ0v5yHHD44XPXCW02ZwVZXPEAXOnO/VBxHPp7m9cLHd5jBHr85QBTncb9ecW8+aFj2wmWpl8G/MTdA/ilxvyW1iz4c6PbC9MeWwfpzNEDhjW+//fegPnha98Jk/pXBGg4yQImfbb92YrM+2Mn3wmbHlMFueS9AXnxrv7Yoxz5YUPjCMLuTwZ9PLUDxmmm/zq/XPjrI+MIOjZLBQ/YqQJJ5Z79Dluk+KjL5wjtIkcE02ShAlHBlv7+h+T5MnPnCXQyQwQHOJkAIE2O/IXcEP3Ii+sINqI7BcRIlQDcCYb8Z7Qo/hoP6wtIEjcFiRyNARNVev3eYET9S4frCSKyLwYq3IUCNVly/7TUXP+Jd+8K2C4rBU14gQLC/Wb+85hs/mLT7wv5YiMGvHh9AR8tWvwkJID/AFvzCbKCGwVMLHkBEUVS/opgjP4Jo/MIu44TBN/YcQOBYUb/E/CY/2t78whQug8FOWRxAttdOvz6YKD8oOP3Cn1uBwc63G0DMYky/S0YqP6Ch/cLelH7BJgobQMccSb+80is/GOD9wp5Je8HlxRpA2mxGv+rkKz98Mv7CQ5d3we90GkA+F0S/CUosP6iP/sKwvnPBuS0aQMPIQb/Phyw/9tT+wtSJb8GFTRpAYX8/v8onKz92Mf/CNd9rwTuDGkBvpD2/7ZspP8qK/8I3MGjBN9AaQJRDO79Cgyc/Atz/wjg/ZMHvNBtAQmE5v4lAJT8fCwDDBBFgwYe4G0ANzDe/E6YiP/w/q0JTrnPCaHLfPm22WD+ao9E+R+OrQmYUc8Jqiec+tQtaP0sa2j75gaxCmHFywoDU7z6E7Vk/hiXiPgYdrULayXHCaWT3PqsXWT96Jek+WrqtQlEbccLIPPw+dpJYP5+j7T4MV65C+Gpwwpze/T7JO1g/8BHvPkX6rkIyvm/CmkP/Pl/CWD/ytPA+oZuvQmETb8K5zP4+C6NZP0Gy8D4NPrBC7HJuwgAd/D48+lo/7LnuPnjhsEJO0m3CCEX4PplmXD+YpOs+2oaxQlU0bcJvaPQ+w7tdP2Z76D41LLJCdZxswuPA7z6fQF8/j5zkPojQskKaCGzCrmfpPsK+YD9nCN8+wXizQpV9a8IrQuM+RwFiPyWH2T5mILRCaPdqwhFw3D74XmM/IWPTPtbCtEJ0cGrCG1LWPpaaZD8h3s0+uWe1QsnwacJs5M8+xLFlP1D3xz7TDLZC0GxpwiwQyT5m+2Y/F7zBPgKytkKe9mjCzFHDPk4kaD99grw+WVO3Qt5/aMJ7nb0+vClpP8pBtz4k87dCRwhowiyAuD5VQGo/r5myPrmTuEJLm2fC6c2zPg1Faz/LUa4+Hja5QjowZ8KsFLA+kAhsP9/mqj5y0blCdc9mwgI+rD5LuWw/slanPqtsukIzZ2bCFIepPmdCbT/F1KQ+cBG7QtQDZsJxqac+G9NtP7oroz7UqLtCoaVlwqL1pT5jTm4/WaShPgZMvEKbR2XCh8SkPnCjbj+UkaA+gOe8QgLsZMIFqaM+bL9uP4qBnz6agL1CEJNkwkPooj5M8m4/ydKePigfvkKKQ2TCxJyiPvpfbz8Pq54+0Le+QkT5Y8JXfqE+nGZvP/yQnT6uTL9CAqdjwkpDoD7dfW8/11+cPmfsv0I0WmPC3rGfPu6Zbz9w2Js+t4PAQtIOY8KtLZ4+GY1vP3xTmj5DGcFCEMliws0gnT75u28/e1eZPouxwULohGLCcbebPmECcD/jBpg+ckXCQnlAYsKcBpo+rFdwP7Bzlj7c4MJCjfhhwsn4mD4uknA/w3mVPpB2w0LAt2HC0K6WPoHYcD+dSZM+ghDEQnh7YcLF15Q+yf5wP82BkT6YocRCIkRhwr10kj6mkXE/eU6PPhI5xULG+2DCQuOQPiL8cT+k3o0+C9TFQmfUYMIn340+vBByP3/mij7xZsZCUZ9gwnXyiz4qY3I/sRSJPoz0xkKic2DCdzuJPs3Lcj/Zf4Y+WpDHQltAYMKOI4c+ND9zPw2LhD4tJMhCqg9gwvDmhT7MiHM/SWSDPsLByEKu3V/CknqDPuAXdD+xIYE+KkzJQheqX8JKq4I+Wod0P3twgD6i4clCmodfwtdQgD6aAHU/HnF8Pjl9ykLmVl/CyS59PtR8dT9oQXk+ygrLQggrX8L1yXk+nSp2P6k2dj7SmctCWf9ewmkxeD4FoHY/LNl0PmYqzEIk2l7ClMVzPrhidz8Fz3A+drPMQtKzXsKe7G8+S8l3P0kpbT6nSs1COZJewh32bT6BM3g/pGVrPhjazULkcF7CmIJqPvS6eD82Mmg+UVjOQo9MXsLSZWQ+zL55Py2MYj4h4c5CpCxewmxCYj52NXo/Z51gPlNkz0JSBV7C1LlcPsgVez8adVs+Qu3PQjjoXcJHM1c+OI57P/IeVj5SctBCDsldwm3MUz7gZXw/vA9TPjP30ELLtl3CIulNPkrUfD/tU00+fHLRQiqiXcLVm0g+uzd9P74oSD5I7dFCM5Vdwk9MRD4rxH0/2wpEPihl0kILe13C46I+PuYVfj9SeT4+eufSQjpqXcLvczg+iZh+Py1yOD6gWNNC3GxdwmTiND4q1X4/VPE0PgLS00KOUl3CzV8vPoSMfj9OTi8+FEPUQu46XcJAsyo+f51+P0OiKj56uNRCKjBdwrgYJj7h3H4/ixcmPiIu1ULQIF3CwX8hPoZPfz/inSE+TpjVQnwbXcLB7h0+PxV/P4j2HT4gB9ZCtwpdwsmVGT5UDX8/upYZPtyB1kLiDF3C39EUPnQcfz/Q0hQ+hvfWQpgAXcJWIRA+6hJ/P4UbED5XXddCtPxcwmLpDT4/L38/sOkNPry310Li8VzClycLPiDEfz83Tgs+kjbYQgvmXMIIPgo+Fw+APwp8Cj71J6JC7NOEwlVhK8C/8Va/BGrevmQ0oUIoBoXCaLQswDosWr9umtW+1EagQss3hcLg8C3AI3xdv9Rszb5oUZ9C52uFwqkIL8B8ZGC/wxvGvlNXnkIvn4XCGvYvwCSQYr8/ub++TlmdQpLWhcIsnzDAiMFkv39mu760XZxC3A6GwllqMcD50GW/a5G1vupem0IkR4bC0ggywFhVZr+95bC+amOaQr6DhsJajDLAf8tmv+kHrb6aW5lCA72GwjYLM8Dxz2a/wiapvtZamEKR/4bC8pkzwHtZZ7+s96S+5FSXQrU1h8JKETTA4pRov9a5ob7+UpZCoGyHwlu/NMBebWm/E6ucvvxDlUK2oofCR0E1wGUha7/BOJm+Xj6UQsbWh8Jv4zXABftsvwbOlL6JL5NC7gmIwnOUNsAe0m6/k+WPvhojkkKTOIjCZzw3wDnkcL8qTYu+wheRQuFliML+5jfArIlyv1V5hr7WC5BCN4+IwgqoOMDYanS/cPuAvk35jkLdu4jCqnc5wDCxdr/zMXa+P++NQvTliMLSLTrAeDV4v+OQa74h4IxC5g+JwgP1OsBqkXm/vb9fvmLMi0I+NInCA8g7wNqeer/nA1O+graKQrFcicLahTzANXp7v899R74MoIlC8ICJwvlMPcAdCXy/C0A7vmWRiELqoonCjis+wImefL97iS2++nuHQnHGicLI9j7AbwF9v8z0IL7mY4ZC/eeJwk6yP8DwIH2/3kUVvo1QhUIhA4rC1l1AwFBpfb/Zogq+hzeEQhIfisIoDEHALpR9v2aV/70uIoNCFj2KwuC6QcC4e32/kLnpvSYOgkLYXIrCnEtCwKhcfb+pnde9b/GAQi93isKq7ELAZ5l9v8Wgw70Gxn9CS5KKwmdEQ8DHon2/k7O4vWeUfUI0q4rCqt5DwN+yfb8efqW94lZ7Qvi/isJtQETAibt9v7tRmb0iNnlCW9yKwuyyRMCLfX2/bvyKvVkJd0LF8orCiAJFwKZmfb/RDIG9P+V0QqEHi8LZWEXAyqF8vwVBbL28rHJChhmLwgiPRcChaXy/3LBevZKBcEKyK4vC/8hFwNsde79owk+9vWFuQr9Fi8IT8EXAsUV6vz2/Rb0kRGxCU1eLwlYPRsBaiHm/3bw9vQYHakIka4vCDiBGwH9zeL8WNjm9ne5nQgeBi8JMQUbAwDp4vwDxML2Pz2VCoJWLwj5FRsCKgXe/fbcvveioY0I+rIvCwkFGwJhJd783fzC9XoZhQja8i8JzSkbA9iZ3v2hQLr1wR19CsNKLwvVYRsAC+na/H7AqvQETXUKB4YvCpjxGwNNtd7+xzTG95f5aQp3wi8IKNUbAuit3v0OVM73p3FhCogCMwlA7RsAsW3e/ThsyvWidVkL2DIzC6C1GwEOpd783gzW9Cm9UQswZjMJIQUbAevB3v8vXML1pSlJCzSmMwq4eRsCifni/C5E5vYgTUEKKPIzC7D9GwJFleL/ZVjG9m/VNQqZOjMKqLEbAmRx5v7VXNr0MwEtCsVmMwr4rRsBq93i/UYQ2va5ySUIua4zC5wxGwN/HeL9DDj69ayZHQkZ6jMKAFUbAaFZ5v14kPL0YDUVC8I2Mwq3fRcAKVXm/421JvfS7QkJepYzCMOJFwEBceb/60Ui9eoBAQjK7jMIrwUXAI+l5v0MzUb0MUD5Cp8eMwga/RcBGFnq/zM1RvcYfPELC34zCWLRFwKYYer/aclS9E8o5QijxjMJ+nkXApxx6v/PbWb10kTdC6PmMwriORcB5vHq/lghevd5dNUKGFY3CDIBFwFA5e79L4mG99w0zQi4vjcJYeUXAIKJ7v4u6Y72FvjBC6D+Nwht9RcC883u/QPBivedxLkKyWY3CZmRFwJPBfL+EcGm9DzAsQrhsjcLdV0XA6bJ8v7uHbL3J7ClCWIeNwvhGRcDYCn2/f+RwveKpJ0IWmY3CeVNFwJxVfb8n62296VUlQpGyjcKiZUXAIu19vzirab2yLyNCo9KNwmM3RcByLX6/hFB1vaDcIELS4I3CgjlFwAt1fr9n63S9E5EeQlX7jcJ8VUXAaLh+vzgQbr36/htCXw2OwrFYRcBurX6/WT5tvajBGUK0JI7CjWJFwEP8fr+n7Gq9RUasQmaBcsIVVDJAHUlrv0ZesD4pW6tCB8FxwgloMUBeMWy/5wW4PrZ4qkJ6AHHCzLswQM99bb8U4r0+/5CpQpZDcMLvCTBA+tRuv7j5wz7opqhCAIBvwr2YL0CYWG+/9LzHPqK7p0JcwW7ClH4vQFzPb7/Bv8g+i9OmQgoFbsKUfy9A4YBvv+2XyD6I66VCTkZtwqORL0BK5G6/NMfHPogFpUKbi2zCvOovQEE/br/fuMQ+dxykQkPUa8KiTzBAwbFtv+FXwT6cMqNCmzxrwoWUMEAnZ22/dBO/PilLokKaiGrCZBQxQNbCbb8zOLs+mGWhQlbRacL9ZzFAQlhtv4t0uD5DfKBCeSRpwmXgMUD7RG2/i620Po+Sn0IMemjCvmAyQKWkbb8m0rA+gqKeQvfOZ8LOwzJATH1tvyGwrT4Ns51CsiRnwhslM0BDy22//MWqPkXJnEJhembCSnozQJqzbb9bGag+itubQgfNZcLdtjNAPgpuv9dVpj6b6ppCPSxlwuHaM0CKj26/lWSlPuwAmkJehGTCPvozQLb8br/Lj6Q+QRCZQmLpY8IHDTRAYptvvw8vpD4xGphCh0NjwmwBNEAgQXC/cMKkPngol0K4qGLCP/8zQFescL9g96Q+1SmWQkkLYsLTAjRAX0Bxv/kLpT7jPZVCUm1hwrXeM0BD5nG/z2OmPupKlEJDzWDC0bgzQFPOcr/44Kc+v1CTQo81YMIhnzNAanBzv7HlqD6lW5JCdpNfwi9zM0CcG3S/4ICqPrthkUJ/9l7CD0ozQPIOdb89H6w+42uQQkJbXsKqFjNArGx1vxnerT6Gdo9ChsZdwkHqMkAiH3a/+YKvPtB8jkJgKl3C4MYyQEu0dr+M1bA+IIuNQgyPXMIwzzJAOWx3v9PSsD5gl4xC/fNbwsCUMkAR3Xe/39SyPmqRi0K4UVvCMo4yQNVAeL+OLbM+JKiKQmLGWsJxgDJACUV4v9aesz4qsIlCdidawqSEMkByk3i/nJizPjzCiEJAiFnCEoEyQK4PeL++hrM+PsKHQjzpWMKjpjJAnRl4vxBZsj6C1YZCLU9Ywgy4MkDrUne/jIWxPnzghUIewlfC+tkyQNngdr9/SrA+jvGEQvAnV8K4EzNAUHR2v1VRrj5h+4NC4JhWwjRYM0DZQna/uxasPhAJg0L4EVbCcpozQBgXdr8C8ak+ZBmCQnOFVcJZ4DNAJMV1v7ahpz7XI4FCovpUws01NEDo8XW/GAClPuxAgEKPb1TC2o80QDAndr9HPKI+Cot+QmLhU8KVzjRAAFJ2vw1RoD6NnXxCCltTwu0yNUDNq3a/7kWdPiLRekK31FLCNoY1QOAHd79exJo+Eu14QhNCUsJy1DVANrd3vzuEmD6XDHdCW8JRwn4zNkCNHHi/aqWVPjAbdUKeQVHCaHk2QNK1eL+bn5M+9EBzQq7FUMIzxTZAxA15v6FWkT5CX3FC90VQwvDuNkCpJnm/mA2QPheUb0J5zk/C6Cw3QH3Ceb97Ro4+5J5tQthNT8KIXjdA2OB5vxq/jD50sWtCqdNOwkWXN0BXxnm/ye6KPtvFaUJpU07CuKk3QBDweb+qZYo+U+1nQvrjTcKs7zdAhNh5vyQsiD6k92VC8n9NwjgTOEDGkXm/PPuGPgYMZEKwEE3CdjI4QHXLeb9aD4Y+gDFiQu+jTMKJUjhAvtV5vywQhT61T2BCDkBMwiR1OEBKoXm/R+yDPptXXkJC1kvCsYE4QGUzeb/caoM+LXdcQjxgS8JWmDhAlWd5v53Cgj5cj1pCtQBLwg6jOEAsqXm/dn2CPuaiWELfsErCpqM4QBSteb+3eYI+m5hWQl9TSsKgvDhAXJZ5v0ergT7RuVRCvupJwjm8OEC9Inq/hdKBPtjAUkLXjEnC2cU4QEnDeb+2bIE+q+hQQo8+ScJS1DhAkqV5v+bwgD5f/k5CvN9IwgzHOEAOtXm/bV+BPoEFTUKahkjCzbM4QFpVer86I4I+yidLQn5KSMKB2zhAkHB6v+bqgD4DKElCj+hHwpzhOEBlcnq/Q7qAPmdSR0JMrUfC9s04QIBuer9JV4E+oiVFQo1PR8LCzzhAcH16v6FMgT6+QENCFgNHwnHOOEABjnq/cluBPmQcpUJeIIPCXxYlwBydSr8BZgS/1zKkQn5fg8JrnCbAMl1Ov/AT/76TTqNCuZ2Dwg8SKMBLOVK/IM31vjRiokJE3oPC+WApwPKiVb9JXO2++3ChQkUehMKWgyrAdklYv7vO5b5Qe6BCpmKEwpNmK8Bo8lq/1yTgvvKHn0IIqITCK2wswBF3XL8o2Ni+upCeQlLthMJxPi3AYGpdv3zX0r50nJ1CGjeFwnz2LcA/TF6/hZrNvnibnEIefYXCh6UuwK2qXr8Da8i+cqCbQm7NhcIVXy/AIZRfvxQjw756oJpC7g+GwisDMMBAK2G/WMW+vgCkmUJQU4bCH9swwNBgYr+Zpri+KpmYQmuVhsIfgjHAw3pkv05XtL5rmJdCL9WGwqhIMsAluWa/rxSvvtaNlkK+E4fC8hozwOz+aL8hbam+CoWVQktNh8Kb4zPAJ39rv14ZpL7wfJRCcYWHwuerNMAAlG2/IZmevjJ0k0JPuYfCkIs1wGbnb7+TaJi+7mOSQhXwh8KkejbApJ1yv3TNkb7mW5FCeiSIwkJNN8CChnS/j8uLvq5OkELaV4jCSTI4wDw7dr//HoW+NjyPQuSFiMKEIDnAopZ3v9oSfL4wJ45ChreIwsb4OcAxuni/Kx9vvngSjUL/5IjCR9s6wD+Beb8bVGG+5gOMQsEPicIR1TvA50R6v7MNUr6m7opCFDyJwjC5PMBLz3q/3wVEvirWiUIoZonCRIw9wB4Te78e8ja+9MKIQnOJicIhTz7Asnx7v2LuKr4MqYdCJ62JwpkTP8AEvXu/JcIevhKThkIr04nC9NY/wLS2e78ikxK+RH6FQpz6icIJfEDA0qp7vwxICL6rX4RCbByKwiwxQcDV9Xu/6yn6vcFPg0LvPorCFJhBwFcJfL87ZO29QDSCQupeisLYRELAUyZ8v8jx172aE4FCanqKwqS4QsCXNny/45DJvcP/f0KgnYrCgTxDwK35e7/UFLm9U819QoS6isJAmkPAiuh7v9Vorb3joXtCFtaKwqYARMCYJ3u/vHOgvcdieUKt7YrCD0VEwHnzer926Je9GC53QusFi8Jpi0TA3aB5v9bQjr1TB3VC7yWLwiS/RMD3yXi/HDKIvYjhckJmPYvCI+5EwKgLeL+4NIK9KZlwQopWi8L4B0XA4ep2v6J6fb3deG5C3nGLwps2RcBotna/pOpxvdpObEK8i4vCjEVFwN73db+15G29Mx9qQtqni8IqSkXAnr11v2OnbL1s72dC4LyLwnJdRcCUl3W/eNtnvWSkZUK/2IvCwnVFwBZpdb/z0GG9lWRjQmfsi8IsXkXAKd11vwvOZ73+Q2FCOACMwptbRcC0j3W/PUtovdwVX0I6FYzCuGZFwEzAdb88qGW90chcQm8mjMKhWkXAJRh2v8fHaL32i1pCtzeMwvByRcAyZ3a/cfRiveBZWEJdTIzCflBFwNr9dr8dsGu98RRWQtRjjMK3bkXA5ed2v3M4ZL2e5VNCsnqMwmBaRcAxsHe/B5VpvSijUUL3iYzCEldFwASNd7/TVGq9QkZPQsefjMKIM0XA72R3v1UAc71J6ExCiLOMwio2RcDiBHi/dKdyvVG/SkL2y4zCdfhEwIYQeL+F8IC9o1xIQmznjMK29UTAQCR4vyxMgb1rEEZCfQKNwpfMRMDXw3i/yIeGvT7NQ0IqE43CEMNEwP4Heb8Cx4e9Y4tBQs4vjcKDskTADhd5vyTWib0dJD9CxkWNwrSURMBxJ3m/yIiNveDYPEKGUo3CiHtEwD3beb/415C9B5Q6Qk5zjcLAY0TAQGp6v63yk70mLzhCzpCNwoJWRMBO4nq/27mVvQTONUJmpY3CZ1NEwDBIe79uOpa9RW0zQtPEjcKZM0TALSR8vx5vmr2oGTFCPNyNwikgRMCzIXy/jticvZDALkIx+43CaQlEwMyPfL+7zp+9/GosQnoRjsIYEETAJuh8v6sVn73TAypC+S+OwrQdRMCWi32/spadvWnLJ0KzU47CKetDwEbQfb/c+KO9BGQlQmRmjsK57EPAByJ+vzXho72lAiNC5oSOwgQHRMAgbH6/87CgvTJaIEKGm47CAApEwFRcfr+yTKC9TA0eQoa2jsLLFUTA97V+v9jvnr0upqVC5O6CwucbJcA/q0q/nVQEv2S8pEIdLoPCcqEmwKJuTr939v6+3tejQnxsg8LQFijAA05SvwKz9b4866JCI62DwntlKcDSuVW/p0Ttvr35oUJG7YPC74cqwPBgWL8uuOW+zAOhQs0xhMK9aivACglbv+gO4L4xEKBCUneEwgRwLMCti1y/asPYvq0Yn0LGvITCsUEtwDp8Xb/zxdK+JiSeQrQGhcI4+S3Ak1pev0mLzb7fIp1C4UyFwtKnLsAXtV6/d13IvoonnEJanYXC2WAvwNeaX78GGMO+SiebQgrghcKVBDDAOC5hv3K7vr6EKppCmCOGwjvcMMBEYGK/vJ24vlEfmULhZYbC4YIxwNN2ZL/kT7S+Rh6YQtOlhsIiSTLA4bJmv64Or75WE5dClOSGwvgaM8D09mi/NWqpvjEKlkJWHofCMOMzwJV2a7/SGaS+uQGVQrJWh8IhqzTA/optv1Ocnr6U+JNCxYqHwp2KNcDk3m+/gW2YvubnkkLBwYfCWXk2wPSVcr981ZG+ed+RQlv2h8KySzfAtH90vx/Wi77L0ZBC8CmIwlswOMC+NHa/qyyFvuy+j0IyWIjCRR45wHKQd7+hM3y+aqmOQhOKiMIg9jnAMbR4v61Gb75ElI1Cy7eIwhDYOsARe3m/woRhvjiFjELK4ojCgtE7wLc+er/9Q1K+f2+LQl4PicIqtTzA/ch6v8BDRL53VopCtDmJwvWHPcB6DHu/jDQ3vspCiUJEXYnCPEo+wJZ2e791Oiu+WiiIQjiBicJRDj/A8LZ7v4wUH77eEYdCh6eJwibRP8AksXu/E+4Svoz8hUJCz4nC33VAwPCle7/1qAi+W92EQljxicKSKkHAhfF7vyH6+r3szINCJRSKwvmQQcCJBXy/30TuvdKwgkJrNIrCZz1CwGUjfL+A3di9lo+BQjlQisLUsELAgzR8v8+Iyr1Qe4BCv3OKwnU0Q8Aj+Hu/LBS6ve3CfkLykIrCp5FDwLjne7/vea69PJZ8QtWsisKp90PAeid7vwqRob3QVXpCuMSKwp47RMCF83q/BhSZve4feEJI3YrCq4FEwJOheb9IBZC92/d1Qpn9isIutUTA38p4vxRtib3L0HNCYxWLwuHjRMBeDXi/kniDvQ2HcULSLovCUP1EwF3sdr/zDIC9VmVvQoNKi8K9K0XAkrh2v2CXdL3qOW1CuGSLwmc6RcDI+nW//aFwvecIa0IsgYvCyj5FwErAdb+bcm+9wtdoQoKWi8L7UUXAOJt1v6Ssar02i2ZCubKLwjFqRcCcbHW/86dkvd9JZEKwxovCflJFwB/hdb8xrWq91ydiQs/ai8LST0XAOJR1v8swa70m+F9CI/CLwt1aRcA6xXW/hJJovYapXUKjAYzCyE5FwM0ddr9/smu9JmtbQjkTjMIyZ0XAbm12v1nZZb1nN1lCNyiMwrdERcB4BHe/b5huvdTwVkL1P4zC+WJFwBnvdr9lHme92L9UQiFXjMLFTkXAHbd3v4FzbL27e1JCsGaMwoBLRcCilHe/RzFtvSIdUELQfIzCByhFwHNsd7992HW9UL1NQtqQjMLWKkXAygx4v9J1db2hkktCmamMwintRMDyF3i/m1aCvSQuSUJUxYzCk+pEwCIseL99rYK9IuBGQqvgjMKUwUTAN8t4v4jlh70Ym0RCqfGMwhq4RMAUD3m/MSOJvWRXQkKUDo3Cq6dEwF4deb+OLou9S+4/QtEkjcIZikTATS15v7jZjr07oT1C0jGNwgJxRMCF4Hm/rSaSvX1aO0LmUo3CXllEwHJver9IPZW9kvM4QqtwjcIzTETAdud6v2ECl72XkDZChIWNwkNJRMBgTHu/sH2XvcstNEI5pY3CjilEwM0ofL/2r5u9VdgxQuK8jcIfFkTAiSV8vxoZnr04fS9CF9yNwnn/Q8Cyk3y/UAyhvZclLUKW8o3CRAZEwB7sfL/sT6C9b7wqQl8RjsLiE0TAXo99vxrRnr0rgihCVTWOwlDhQ8Be1H2/fDSlvbsYJkJASI7C+eJDwGkmfr8UGqW9V7UjQgpnjsJX/UPALnB+v1jnob3GCiFC7H2OwjwARMCyYH6/AYahveq7HkIvmY7C+wtEwMe6fr/6KqC9gq3OwqJszsE50hhAz8I6v7k+Lz8cZM/CsCrMwb1VGUAa2T+/xi4vP+EI0MJa7snBsvYZQLd7Qr/Bqi0/7q3QwrTCx8H4PxpAMJRDvzXuLD/HVtHC4I7FwaA5GkCR8kK/9MksP9H30cKIbsPBwyAaQMGyQb+esyw/epjSwitVwcGT6hlAGeo/v/fdLD+COdPCuC+/wZhkGUC3Zj2/K/4tP7vZ08I/Db3BGRwZQFTJOr/4GS4/rHPUwpHmusHDgxhA/jE4vyZzLz9uDdXCKL+4wejkF0B2Aja/TgsxP5ik1cKkhrbBI2wXQPYZNL9CJTI/1jfWwq1mtMHH1xZATucxv1qLMz9+ydbCiDqywZKMFkAdTjC/0Qs0P+Bm18KjEbDB0j4WQFsTL79JvDQ/ePHXwjXmrcEc8BVAVsotv2hpNT9vgNjCisOrwciqFUDFryy/7gM2P1gG2cImrqnB0HcVQJAIK78SGTY/F4vZwtp1p8ErRBVAHHApv+Y1Nj/fDtrCCmSlwYTyFEC73Se/5sk2P7WM2sLxOaPBQcgUQPZ6Jr9u1jY/cf/awtMmocE2mhRAS8Akv6rKNj82hNvCfh+fwfOBFECUKyO/UHk2P7Tz28K7G53B8VkUQEuHIb+rXTY/P2zcwjD/msFIhBRA3UIgvxssNT/c0NzCKQCZwQqAFEDRsh6/VY40P4VG3cJACJfBdYUUQJGdHb/MADQ/Ur7dwoMilcG3fxRAEp0cvzGnMz86IN7CExeTwRqWFEBq6xu/KwQzPzKK3sIpBpHB0ZQUQLAHHL9eFTM/iu3ewmAdj8F3exRArFIbv5MnMz/cWN/CKC2Nwc9vFECPrRu/xnszP3C838IpKYvBulkUQNjnG7/G6TM/8SLgwnw+icF+TxRAw58cv2dhND/5g+DCl1yHwd8KFEBtIB2/sqE1P+jr4MKYU4XB9s8TQAnVHb+P1DY/fE7hwsV0g8H+fxNAeaodv/f2Nz9wsOHCE3+BwfA0E0BFdB6/FnQ5PzYD4sIABX/BA7cSQEmbHr/1bzs/XGHiwlwSe8GuORJA1Kwev75hPT9muuLCazR3wW/CEUCSrR6/xDU/P7IE48JBSXPBtBIRQM0DHr9ymkE/lFXjwmS0b8H0ohBA+5Udv9UfQz/VtePC+81rwfEeEEASWx2/PA5FP3AI5MKp22fBbIIPQJXYHL8EPUc/d1TkwqsVZME7Ag9AlvIbv6LLSD84peTCPFxgwQ6HDkD8JBu/9lFKP47t5MJtmFzBaukNQIGFGr/+eEw/6DjlwkzpWME5kA1AD9QZv4SFTT9Ye+XC+upUwTMTDUAZ1Ri/XPpOP0XK5cK4FlHBO7AMQOkaGL8iKVA/pwPmwvzvTME07wtAYNMWv4iJUj/MTObCOxBJwVaWC0ARDRa/0YhTP5SV5sKTGUXBNzALQLgQFb/koFQ/FNHmwq0dQcHOuQpA6P4Tv8fuVT9hJOfCUto8wb4mCkB6gRK/V3ZXP1pe58KmFDnBFqcJQG/9EL81qlg/LZTnwrAtNcHIFwlAn+4Ov/TPWT/c5+fCntswwfOhCEB4hA2/MeZaP+gg6MJFqSzBx/oHQLhJC7+ZTlw/RlvowumtKMFeZQdAVn4JvxaqXT/GkujCH40kwdXBBkD7JAe/gutePwDf6MISjyDBukUGQApTBb8h118/lR3pwghUHMGBrQVACFEDvyEWYT/+Y+nCikMYwZ05BUCQ9AC/DophP1Cj6cL/KhTBeIMGQLyE/L4HpWI/WdrpwqfZD8ENuQVAAgT4vjBPYz9qGurC2eYLwX4pBUBHvfS+edNjP1FC6sIDlwfB/CoEQEBb774UVWQ/taTqwlyXA8FYnANAic7svmk0ZD8u2OrC8jj+wK0kA0D45+m+jbdkPy8d68IjVfXAF0UCQIp05b5f2GQ/rz3rwu0t7cDJYgFATlXhvtCiZD/2iuvCSfTkwBAdAUDe9N++x6ZkPzSz68IsONzAWbQAQMEI3r72jWQ/Burrwg/C08C7PABAeX7cvsPsYz/wKuzCDL/LwKF8/z/oz9q+lU1jPxRw7MK3FcTAQT/+P8ox2b4sJWI/PJvswsPWu8BZ1f0/SanYvoHBYT8UtezC1hizwPgy/D/SAta+W6NgP6qfrULEZnPCYTI2QE3Tb7+HOJM+k66sQr24csL8MzVAwuNwv29umz4AxqtC/gpywkB0NEDNTXK/dtqhPvjXqkJCYXHCG7UzQIy6c79LTqg+UuepQhKxcML3OzNAqE50v51OrD6C9ahC2wVwwuIbM0DIz3S/on2tPrUGqEL9XG/CTRQzQAd5dL+onK0+BRinQqSxbsIbIDNAAMhzv0gArT5nK6ZCgApuwrtwM0C/BnO/LjaqPjQ7pUI8Zm3CNMwzQKVTcr96HKc+nUqkQjbibMJkBTRAdeJxv0ItpT4nXKNCbUBswlt3NEDuHHK/zbGhPqFvokLKm2vC5bw0QIyUcb/UWp8+MX+hQtsAa8LxJTVAhWpxv2kInD4Mj6BC+GdqwjqTNUC/tHG/YbmYPjGYn0JVzmnC6+M1QByBcb+/J5Y+CaKeQho1acKOMzZAE8dxv5fDkz6JsZ1CSptowg51NkAqpnG/c7GRPni9nEIP/mfC3Z42QML2cb/VfJA+tMWbQghtZ8KxrjZA/Hdyv3gkkD7V1ZpCrdRmwvy7NkD21XK/3dWPPr/emUK2SGbCdLs2QAJjc7/DApA+q+KYQnKxZcJvnDZAYPBzv+oikT6K6pdCBiVlwvSGNkBFQnS/JuaRPj3mlkLLlWTC/nY2QMu3dL/4h5I+/POVQsYFZMJaQDZAa0R1v1lmlD4U+5RCZnNjwn4INkABB3a/2F+WPk37k0Lt6GLCvtw1QDqEdr+k5Jc+dgCTQqpTYsKsoTVA/g93v+/pmT4rAZJC28NhwjRpNUDi33e/pPCbPtAFkUKPNWHC0ig1QBAkeL81DZ4+KAuQQnatYMKp8DRA7rV4v88AoD4KDI9CoB1gwvTANEB/Nnm/u6uhPnYVjkLGjl/CFMA0QIHWeb8H5qE+wRyNQg4AX8LsfDRALzR6v/QkpD4oEoxCSWpewnFvNEBXinq/aq6kPh8ki0JP6l3C9Vk0QM+Der/RWqU+bieKQvdXXcL+WTRAVMd6v5RwpT4oNYlClsRcwh5TNECGOnq/cnqlPuAwiEIaMVzCfHQ0QNE7er/ua6Q+gD+HQiiiW8JwgjRAm2h5vzy2oz7RRoZC/yBbwnSiNEDv63i/s4qiPjZUhUK2kVrCltk0QFZ2eL8jp6A+5lmEQogNWsKQGzVAqTJ4v8l8nj4+ZINCxZFZwqZaNUBoA3i/AnGcPuRwgkLwD1nCtp81QDyrd782KZo+PHiBQm6QWMI78zVAjM93v7yTlz5ckYBCAA9YwplJNkB+AXi/fuuUPiYlf0IEjFfCCoY2QEwjeL8bD5M+yC99QjQPV8L25jZAw4B4v3AekD6bXXtCpJFWwtc2N0DY1Hi/vrONPupzeUI4CVbCx383QHx2eb/klYs+G4x3QqKRVcI12TdAp9l5vz7hiD7WlHVC7BlVwq4YOEBIa3q/UwmHPtW0c0JbplTCc2E4QDzDer+D1oQ+GM1xQuAvVMJ1gjhAr9R6vz/Rgz7K+29C6MBTwtG6OECFYHu/Vy+CPrUBbkJ6SFPCqeU4QFZye7+r2oA+kw1sQsjWUsJgGTlAZUp7vy9gfj4wG2pCSl9SwvsjOUBoa3u//MV9PrA9aELB91HCK2M5QJJCe79zuHk+IUJmQm+bUcKefTlAJPJ6v5jndz7PUGRC4jNRwm2WOUAzHnu/JW52PjJxYkIGzlDCG605QH4Pe79T+nQ+1IpgQoZxUMKIxjlANsh6v2k/cz6ZjF5CdA5QwkDMOUCRTHq/SKhyPiunXELgnk/CV9k5QPdxer8L6HE+2LpaQv5GT8IV3TlAqqJ6vzzDcT6wx1hCNv1OwpnUOUA8mnq/c0dyPsa5VkIWpU7Cj+M5QOJ0er9jRXE+2NRUQo5DTsIb3jlAKfZ6v6HacT4A11JC7utNwjLgOUAtiXq/F4VxPjz4UEJUo03CUOg5QCZmer8X8nA+hwhPQjhKTcJJ1DlAQG16v+w2cj4WCk1CHPdMwt67OUBkAXu/NAZ0PromS0KHwEzCId45QJQZe78q63E+riFJQrFjTMJq3zlAuxl7v5LWcT6rRUdCxCxMwnfHOUCAEXu/i1NzPuMRRUIY1EvCFcU5QKoce788f3M+WCdDQvKMS8JxwTlAtSl7v/6/cz4HMdrCzmeGwTObO79AIS8/CXIlv+5x2cJaHYnBez06v8h4Mz+vxyW/Ga7YwmXDi8GiADy/b7M2P429KL+c89fC42GOwW72Pr+y9Dg/eIosv2ND18KlEZHBI29Bv1XOOj/6uy+/PI/WwifQk8HO6UK/Xs07P46dMb/H2tXCmp2WwTSWRL+ryjs/2Eszv0Yv1cJAfZnB/nJGv8M5Oz8f8jS/nnrUwkxsnMGTw0m/2D85PxB7N797yNPCcGKfwYPkTL+BkDY//4M5v6YU08J8T6LBkeJPv6FVMz8IKDu/fmbSwv5PpcF3ZVO/0P4vP9I9Pb8ks9HCbG2owYvXVL93aC0/Eo49vwUC0cJVhavBHjFZv+u2KT8zQ0C/lFXQwouNrsFvN1u/HlUnP7E0Qb9wnM/CQ86xwSv1Xr+UDiQ/mW5Dv4LqzsKM5bTBdK9hvwkrID+WU0S/sjDOwnoOuMEz4WS/bp0cP4vORb+Lg83Cciq7we1rZ78qsBg/82xGv9nWzMLHdb7B91BqvxVBFT/ymEe/DCfMwiJ2wcHHQmy/7pYSP0UwSL+FdsvC1brEwSP6br9McRA/F8dJv9nAysLZ4MfBAFlwvwx4Dz82oUq/oxXKwkwOy8Gtj3G/0TgPPxSyS7+JZcnCTyTOwU9Kcr/WAA8/9UxMv7a1yMLFOtHB9G9yvyLQDz813Ey/9wzIwoJD1MEE1XK/ZQIQP7BZTb/4XcfCAkfXwWNuc79X9hA/xm5Ov4muxsJXUdrB2Qx0v8qIET97V0+/xgnGwqg+3cHRNnO/ooQSP5MET7+LWsXCgDzgwRNUcr/y2hM/rNJOvyarxMI8FOPBM8Zwv8hNFT+jA06/qADEwiT+5cHvO2+/zBUWP2jgTL+oVsPCarbowXPkbb/sJRc/exNMv9HAwsLhiOvBucRsv0nfFz/pUUu/2A7CwkZn7sGBt2q/9PEYPyDQSb+ocsHCmyHxwd1yaL8xJxo/kidIv3jYwML7yPPB+IBmv0R1Gj9hYEa/2TfAwraL9sFgAGW/+bQbP/B9Rb8CmL/C9TP5wVh4Yr8S0xs/tQxDv8r7vsKCxvvByi1gvyNdHD/wDEG/5VW+wm16/sGS6l6/QtccP7wIQL+Wxb3CNX4AwrGAXb9goBw/6ow+v2wivcKMtQHC+axbvxyOHD/8uzy/FJe8wiTuAsIOrFq/1YQbPyNJO7+88rvCDSIEwt9mWL/TFhs/6eQ4v65Yu8KjVgXC5g1Yv0WwGT8U7je/pM26woh5BsJS7Fa/9nQYPzdLNr+BNrrCKLwHwjAzV7+jkBY/j7Y1v/ymucJz1QjCfCpXv5/bFD/76jS/zBq5wkfuCcLjnVa/fE4TP7a0M7+JhrjCeSULwlnHVr/XPBE/qe8yvyn7t8KaTgzCytlXv18zDz/hBzO/QX63whxlDcL93Vi/r5YNP7ZBM7/p9rbC2G8OwvetWr8rrAo/L6Azv/RvtsJNiA/Ckx1bv7xSCT+CajO/2+K1wkmpEMK3Q12/Hi0IP6DiNL8UWbXCU8IRwvbbXr+OnQY/mKQ1vzDWtMIn3BLCeRNhv/t0BT/HKDe/TlC0wlT9E8IEsWO/hRUEP7rwOL/2xbPCgAsVwkVAZb8eCAQ/L146vwJQs8KEJxbCTgxovwz0Aj9WdTy/NeCywsM2F8LcRWq/go4CPwZZPr80ZbLClE4Ywkogbb+aLAI/N9dAv9LiscIIdhnCsC1wv3g6Aj/IwEO/pW2xwqKRGsKyaHK/BkECP7viRb8e7rDCcbUbwoUjdr8eMQI/P2tJv6R3sMJq1RzCjbJ4v9UEAj/9yUu/oPSvwmDsHcKpgXq/SuACPyX+Tb8Sca/CISEfwv2KfL/yGQM/0RdQvxL8rsIzSSDCSzt/v9emAj/ZeVK/8oauwgVoIcLXc4C/wK4CP08iVL+eA67CEJIiwoAZgb+lMwM/pLFVv310rcKksiPCNF2Bv5fgAz/9lla/aAGtwkTcJMJIloG/rdoDP7cEV796oazCowImws0Wgr9e7AM/Zw1Yv8wGrMJVNifCuFaCv6blAz+eiFi/cJurwjhfKMIfQYK/0NYDP21VWL9YDavC43wpwjXygb+iKAQ/S+ZXv6ieqsKawSrCa+CBv6AEBD/4rle/XGelwrZemELFgxxARi9Lv6iWJj9OrqXCy6CYQjA1G0BfYUi/Ot0qP+PqpcKK3phCS0UaQOY0Rb8teC0/Zi6mwkobmUKqhBlA0vxBv+tGLz/Mc6bCMFqZQuDRGEB/kD6/gMEwP+q5psKal5lCzisYQH6hO79YMzI/vfqmwhTTmULxwRdAqzE5vxThMj9WP6fCVhCaQsFHF0Dc1Ta/yNQzP06Fp8LHTppCPhQXQF1WNb9QBTQ/C8anwuyKmkI41BZAs1c0v+ybND88CqjC28SaQu2sFkDhKTS/3iU1PypEqMIrA5tCu6kWQJRWNL81RTU/wICowmQ9m0KsrhZAzI40v89INT95vqjCaXmbQnfaFkAnODW/1980PzcCqcJHtJtCCCIXQB/PNb/j/zM/Czepwoztm0JqNhdAAE82v+TiMz8sdanCMiOcQkljF0A7qza/KFUzP6GpqcLZXZxC6oUXQNyYNr84wzI/6dqpwiGVnEIpihdAzjU2v+aJMj9ZFqrCic6cQi2WF0APqjW/DyEyP2dIqsI3Bp1Cp6UXQAz6NL/bmzE/wnmqwtBAnUJ5thdAuj80v1YNMT98sarCfnSdQpzgF0ACTDO/nAMwP7bfqsI2qp1CQxAYQMV2Mr9P8S4/8g6rwl3inULZRRhAfdcxvzveLT/sOKvCyheeQgaEGEBbMzG/bKgsPyhrq8KMTZ5CicMYQP1SML+aVis/oZ2rwhZ+nkIkCRlAHa4vv90FKj+NyqvCRLWeQjRDGUAM5i6/cdUoP0oBrMKC6J5C6HkZQA6mLr836Cc/dS2swscen0JcfBlAPLQtv3CAJz/lWazCOU6fQgGEGUA02Cy/KQ0nP9SIrMKuiZ9Cq2oZQBnSLL+SbCc/MLqswrm7n0JkPhlA+4Esvzb4Jz/a5KzCHuufQnr2GECO8iu/m9UoP7sYrcJeIKBC9qcYQLgBLL9ICyo/JkqtwgFToELqORhAd8Qrv1SdKz/Ifa3CKYqgQh2TF0D4fiu/hgkuP3WvrcKMvqBC/+4WQCaeK7+vljA/U+Gtwi32oEI3VhZAw8grvx7/Mj9MEa7CuCyhQsFzFUAHvSu/EXY2P/I9rsK7YaFCcqsUQDvXK7+umjk/6GOuwiSQoUJp/hNAymIrvzwYPD9im67CCMihQq9WE0DFuCu/oNw+P7DHrsKq/qFC264SQKL9K7/OnUE/rO2uwgc4okK94RFAOocrv+ejRD9AIq/C62iiQvFKEUDXWSu/SvJGP3BDr8LOn6JCKscQQHadK79sKkk/WXCvwnbVokLYTRBAanorv6QJSz8SkK/CxQujQpzmD0DrRiu/KphMPwfBr8KNR6NCBocPQIrBKr+b4E0/IuGvwmGCo0J0Pg9AmNUqv+kUTz91/q/CprWjQtPhDkCaQiq/NUtQPzUusMKO7qNCm7wOQGwHKr/Sx1A/KkawwmQlpEJipA5AsiEpvw67UD8qdLDCYGCkQrFyDkBYeii/jTVRPyWUsMJklqRCZYwOQFsbKL/snFA/Bqywwu/OpEJEeg5AfHInv1yUUD9t5bDCpAilQuWoDkAEFSe/8KZPP4AFscIpRaVC8KQOQO4tJr+2RU8/9Sqxwv57pULLyQ5A6vclv36UTj9BRbHCiLKlQiXsDkCMYCW/R75NPwx3scJT56VCpkQPQDV9Jb/WY0w/WKGxwmEkpkJ8iA9AknIlvxNLSz8Fx7HC8FSmQhMBEED8UCW/dlJJP6PwscKgjKZCdVYQQLEPJb+x2kc/qBuywqbJpkKYuRBAb7Qkv7IgRj+PSbLCe/emQjc7EUAJMSW/N1NEPxZissJCK6dCsIkRQMUHJb8iBkM/g6Kywq5jp0IoNBJAjlolv4+EQD9MzbLCrpunQkyKEkBEACW/PwU/P/j6ssKg2adC1fkSQPmDJL+tEz0/LiOzwpgOqEIifBNAvJ8kv1EeOz9AX7PC0UOoQnnmE0CfbyS/G2c5PwGDs8JRdKhCxmMUQKwKJL9GUDc/P7KzwnqvqELI5RRASusjv7RHNT8P3bPC9t2oQltpFUC1mSO/rSUzP1IQtMKGB6lCGwkWQPY5I799kzA/B0S0wqY4qULAeBZAwDQjv2bjLj+Aa7TCZGqpQg73FkDOdCK/Ua8sP8id7MKR6QHBU2c6v8u7MD9S5SS/1Onrwo14BsEZ7zi/NKM1P9NTJb/UL+vCJOQKwdIFOr8onDk/7+Ynv92B6sJ2PA/By9E7v7SMPD/Hziq/7dvpwnutE8HdGD2/BvI+P2kCLb+DNOnCqjQYwfl+Pb8Qh0A/FQYuvyqM6MLc2RzBj8U9vzQxQT9cjy6/vOznwkCRIcFUTD6/KWpBP7UtL79SRefC32EmwSE6QL9WPEA/yKowvymf5sIaOSvBWRFCvwJaPj/3xzG/S/flwiLuL8EBoUO/wc87P9lWMr9cVOXCk8I0wU39Rb/uCTk/V5czv46t5MK50TnB7qRGvx7INj+eUzO/RQjkwuTNPsEp/Um/b3IzP9pKNb8LZ+PCLatDwRqiS785JTE/8fc1vxG34sJU5kjBxE5Ov4gdLj89WDe/Dg/iwrbUTcHodlC/Z3wqP6fsN7+pXOHCc9dSwfbkUr+GbCc/+f04vxC44MIR1VfBz7dUv/AIJD+9Szm/SRLgwnEHXcGz+la/v00hPwpMOr/GZt/CuL5hwaIrWL8qYB8/sZg6vwm93sKE3WbB4Bhav1ETHj9j4zu/tgnewmrQa8H6vlq/DP0dP1V7PL+PYd3Ce8Zwwa5QW7/gkx4/gk49vyCz3MKmj3XBQTdbvylFHz+Khj2/5wTcwjVbesHL41q/itwgP2LuPb8XXdvCIwF/wbqpWr/C9iE/vTU+v6Ct2sJo0oHB6JpavxO5Iz8B9D6/pADawvUthMEkyVq/IPEkP9yvP7+mXdnCGGWGwWpTWb8yfiY/0fA+vy6t2MKGtYjBNiNYv7BfKD8Zmz6/Gv7XwtDdisHXMFa/7SMqP5d0Pb93UtfCRxONwVJPVL+eMSs/Ggw8v2qs1sLSIo/BNbJSvypiLD8c9jq/4BDWwqlLkcHHY1G/sEotPyUOOr9vYNXCPX2TwaPwTr+CXS4/YRQ4v7LF1MJYkpXB0X1Mv26wLz9tNTa/qCrUwjSSl8HL3km/JDMwP67SM78Ci9PCCrCZwTfuR78igTE/MHEyv6Ln0sKNspvB3NhEv1W5MT9FfC+/8krSwjGjncGA5kG/FGYyPxbbLL+GpdHCuMGfwT7cP79cNDM/xCsrv7YS0cIKq6HBhbc9v58xMz8DESm/EHPQwvh9o8HDGju/eWUzP+qXJr+m4c/CSF2lwc4rOb/eqzI/328kv4g+z8IoM6fBlAk2v0qJMj9JWiG/cKPOwoUNqcEGvjS/3GgxPyCwH78eF87CesOqwbPFMr/rVzA/kmgdv1l+zcKMu6zBL0MyvyGVLj8uSBy/benMwlFursHqXjG/peIsPzfSGr90XszCRhWwwQkzML8uXys/pisZv/7Ey8LaBbLBXeQvv3RPKT/IJBi/bzTLwkPgs8HSdjC/D0UnP4jyF7+Zs8rCto61wfhgMb+0dyU/oyYYv84nysKkKLfB/JYyvwqIIj/6NRi/xJ/JwufYuMHWQzO/xQghPxRKGL/ZC8nCtae6wbCaNb95gx8/puMZv66AyMIJY7zBaF03v22qHT8W0xq/kvzHwvwgvsFQ4zm//04cPyijHL+/dMfCne2/wTLnPL/AnBo/KMQev5XoxsLclsHBPy4/v99EGj8ovCC/yGnGwlhfw8HQY0K/J/UYP3MvI7+69cXC1g3FwWhjRb/uVxg/Dbglv/ByxcKJzsbBHPZIv/q3Fz/uySi/9+3EwqCmyMGLyEy/6nUXP4hBLL82dMTC8XDKwRHyT7/4VRc/Ki0vv2fxw8K2OszB2FVUvzIvFz9yQzO//nXDwgAQzsGayle/SdYWPx1mNr9T6sLCb87PwYw4Wr/3jRc/UQw5vzNowsI0uNHB5jVdv2uWFz//8Du/LefBwqqT08EHbGC/KTMXP2jePr9bb8HCpFfVwa23Yr/DPBc/Nx5Bvw/kwMJMMNfB8p1kv62VFz9fI0O/u03AwgAD2cHp62W/8xwYP8GrRL/z0r/CruDawfbSZr+zARg/soFFv7pov8IrrdzBF05ov0r0Fz+a8Ea/nMi+wmCe3sFXWGm/RtgXP7npR79eT77C2XjgwRB9ab/y0xc/4wtIvz29vcK1JuLB9zZpvx7fFz8bzEe/v0e9wqsu5ME9VGm/bbUXP5fUR7//zN7CBvM4QqfIQ7/67jE/u4UuvxKh3sIyuThCPNJKv8gDMz9Y8TW/dHvewp2FOEJGgk+/0lI0P68zO7/9Y97C5lo4Qh/IUL9JpTU/DQ89vwJK3sLgNjhC9VhPv5ZINz/VTjy/SjHewroUOEId6ku/KCA5PzWbOb+RF97C+O03QjVUR79wKDs/os41v3wA3sLEzTdC+EZCvxGbPT/+sTG/JPHdwraxN0Ipaj2/n4c/PwuOLb8a3d3CvY83QtoeOb/dM0E/KeIpvyPI3cIshTdCH/Mzvw4vQj/MEyW/87XdwmxqN0L2WC+/mcpCP+C3IL9KoN3CFkY3Qk/CKr/wT0M/EVwcv0yP3cJeJzdCt2onvyc8Qz+MChm/GoPdwnwLN0IcnCS/sVFDP9pQFr/kbN3CaOI2Qvl5Ir8Af0I/lfUTv8ha3cIhwzZCjdEgv8ESQj8iNRK/GT/dwmqqNkL8/h+/QuVBP+lZEb9eL93CLoE2QkajHr84zkE/ZQEQv8Yh3cIhYzZCQ7MevxrxQT8OHBC/YAjdwsBHNkIkyh2/SrNCPx54D7+w9dzCAi82Qvo/Hr/JM0M/axMQv3Xf3MIlBzZChKEev/irRD+Y6hC/ksTcwmbnNUKRbh+/qfhFP6UdEr8TrNzCK8o1QphoH7/Eb0c/g5ASv7SQ3MJ+qzVCETUgv4PVSD8UzRO/6nfcwpKcNUKzKSC/R8lKP4RjFL/aWdzCEHk1QlO5IL+xgEw/YoAVv3093ML4WjVCkpUhvyIMTj+X3Ba/yizcwm5FNUKoCyG/3k9PPzy8Fr+qCNzC+yY1QqgQIb90JFE/tVkXv4Tu28JeDzVC4Vsgv3doUj8aDRe/IcvbwncBNUJCxx+/aItTP27VFr/CutvC3Og0QvXOHr/BoFQ/IzQWvzqU28LKzDRCTwsev0/mVT+S1hW/pnPbwqetNELiRxy/9f9WP2FnFL/2XNvCd5g0QoXYGr9NWVg/9l8TvwtA28KGijRCuUUYv9piWj/pZRG/DCvbwphtNEKaXxa/8dZbP0nqD792DNvCt100Qip2E7/dWl0/VWoNv9bt2sKqTTRCeBQQvyJOXz9pjgq/Ks/awp4kNEJeLA2/tMdhPxpQCL/+sNrCuAs0QlunCr9sdmM/EDkGv8aj2sLx+TNClmQHvzKqZT82ggO/tIPawtjeM0L0WAS/OU5nP8fYAL98Z9rCJcgzQjH6AL8tM2k/GtH7vptJ2sIaqTNCk+78vlviaj9Cjve+jjHawryXM0JVr/W+YANsPy+78L7IFNrCqXUzQvVz8b7ddW0/7x7tvlnv2cIIYDNCbKvrvjqhbT/8T+e+LufZwmROM0KaAue+IrZuPzcT476yvtnCJyUzQgAs5L7eUm8/q3fgvqaf2cKd+zJCDzjhvseWbz+qld2+EInZwibkMkJzfN6+patvP0XY2r7bbtnCzMgyQmjp277dSm8/TRHYvtZa2cKksDJC97zbvuIKbz9hyNe+djDZwpWJMkL0htu+yEpuPwc+1751FNnCxGwyQg3k2r43gW0/0EHWvrwF2cLQSTJCKHzbvsq5bD+HhNa+ZvHYwm4tMkLPXty+mstrP+YA177M2NjC3xMyQvpX3b6xwmo/YYfXvuaw2MLa7TFCbt3evgMjaT/0Vti+RqHYwtXHMUJHMuG+Yz1oP65G2r7ig9jCG6sxQukm5L4iM2Y/GFDcvoJ02MIPjjFCc6fnvqu8ZD96I9++6mbYwu5dMUJJR+u+MeFiP8fi4b7HR9jCP1UxQseJ777w3WE/BaXlvq8y2MKsGjFCqpXzvlnmXz8RuOi+4RTYwsr2MEJxovW+uxJeP17c6b61/9fC39QwQliw+r4k0Vs/xr3tvpjp18J2rTBChqj+vpQrWj/z0/C+6OHXwsacMEKjFwG/KoVYPw518755x9fCrIQwQlu4Ar/4TlY/Y4L1vqG018IDXDBC1bkEvwFcVD9favi+4qnXwkIrMEK+PQa/boFSP9tj+r7mldfCIBAwQp1dCL8zZ1A/1GX9vtyB18JR5C9CTwMKv+OVTj9em/++VW/Xwte0L0J20gu/6mJNP4o6Ab9lbtfCQKovQp9IDb8WSUs/KQ4CvwBW18KYeC9Cl0kOv1N0Sj+YygK/Cob6wtrQyr+sCj2/mK0xP9rRJ7881vnCuEXsv9XdPL9nojY/rpApv2gg+cLzSQbAWks+v5MROz97syy/PHj4wvMfFsAvnj+/3G0+P3pXL78M1/fCljwmwK8IQL/2NkE/ctswv5A098ITpTbAjFA/v3M3Qz/k6TC/mJL2wgOAR8AsAD6/L3BEP+kNML+U9/XCYXlYwF67PL/8RkU/uRYvvyBW9cKm0mnAd9w8v3G2RD+GAC+/KLT0wsgre8CM1zy/5WlDP/16Lr/MD/TCl8CFwJSlPL+GOUE/A3Atvxpw88JjRY7A2F89v+C+Pj/XNS2/Esvywv88l8CC0jy/zpE8P77QK784KPLCgQWgwO/dPr8wWzk/OZosvyiI8cLMkajApc8/v3AFNz+Xnyy/atfwwpO1scBnG0G/syk0P4LGLL+iLfDCSly6wDZtQr/84zA/cMUsv1R378LLC8PA7ddDvx5fLj/oJS2/bs7uwjrXy8BUlES/iMUrPx/RLL+aI+7CcMnUwOwERr+p4Sk/r3ItvxRv7cIQ59zAJDFGv3LvKD/UOi2/WL7swsKq5cA8Fke/6K4oP0j+Lb9IAOzC+zLuwJm7Rr+QwCk/ZxYuv4ZK68K6svbA8V5Gv2htKz9/ay6/vI/qwiXb/sBzLEW/CkItP0T/Lb9C0enCXX8DwUsCRL+J6y8/7e4tv3wb6cJQXgfBqaxCv3VIMj8cki2/KFrowkI/C8E1f0G/5jA1P2SRLb98m+fCLjkPwZ6jQL8WWjc/BpMtv6Pn5sJK4xLBh/o9v/DLOT885Cu/ICLmwne1FsGXuDu/jKk8P77AKr/TXuXC2kUawbuYOL+8KD8/fJYov+yf5MLX1R3BabM1v23dQD9oVia/+uTjwpgyIcGn5DK/rq5CP7Y0JL8RLePC6rokwZBiML8ULkQ/MT8iv19n4sICRijBP6ksv9i/RT8vGB+/irXhwoakK8GuHCm/pKZHPxQ5HL+//uDCNNkuwTXzJL/p0kg/j3wYv6ZF4MLCQTLB1Zchv76qSj9GxBW/+oXfwshqNcEzBx2/LnZLP4OBEb+RzN7Cam84wYpZGL/Q9Ew/rlsNv+wJ3sIW6jvB7aMUv8a0Tj85PAq/cFfdwpvsPsHytRC/BZZPP52hBr+mm9zCdsBBwWB7DL+6v1A/DtECv3Pj28JlrETB/pcIv0MuUT9wRP6+xCHbwoONR8GWzgO/tFtSPy2Q9b5kZNrCb2xKwbmMAL96dlI/QEjvvj+02cJ0A03BFmj5vs3GUj+T++e+YPjYwsEhUMGL8PS+z0pSP2Ft476iO9jC6sdSweZL775i7lE/UNHdvpCO18IOQFXBRaTpvg7FUT/qT9i+hszWwqljWMFL6+W+CA5RP2hr1L5YE9bCAE9bwVXC476MW1A/eArSvihr1cJx2V3ByALjvqG9Tz+zDNG+MrTUwoE/YMHj8OG+8/hNPwM9z77yA9TCb89iwbZZ4b6mok0/J4fOvldE08IwnWXBHKrjvsD2TD+RatC+mY3SwmZFaMFt5OS+Fu5LP8Qb0b704dHCIPxqwV0o6L4iaks/MfXTvpgx0cLFwG3B6lTsvjpxSj8Ac9e+VXvQwpg0cMGbMvC+IZJKP6cq277Cys/C9vtywUDS9L4M3Ek/xjPfvlAqz8KwfHXBhO35vmvAST8E/uO+FHjOwi0zeMHkSQC/JmZJP74i6r6Uxs3CKfV6waa6A79yd0k/gbvwvlIdzcKlwH3BKPoGv6qfST+lCPe+1GzMwrUtgMHrNQu/Z9RJPzVJ/74VxcvCRp2Bway+Dr8ms0k/sgQDv7AHy8KO6ILB8kARv12WSj8HtQW/qF3KwiZehME8uhS/bqVKP3wbCb/rqsnCS8WFwQEEGL9OlEo/Rk0Mv8gHycK2GIfBb4YavyrGSj8B0w6//kjIwut7iMGD4xy/b1RLP5pTEb/BhsfCcuKJwSa8Hr828ks/JFkTv07bxsLjUYvBQhsgvzwGTD+zuxS/7T3GwpujjMFqGCK/JiFMP3K+Fr/Ac8XCjCyOwUbCI7+oJUw/L2gYvwzHxMLwk4/BtGIkv/ZwTD9LIRm/3gfEwszCkMHHtCS/wIlMP4t7Gb94YsPCTVSSwb8/Jb/tyUw/8Rsavxd35EICpkXCXXEvwJvFWr8xocC+wWnjQiYaRsJwNDDAKItgv30Fvb7dXeJCqo9Gwq7/MMCx62W/GOK4vqJJ4UIFCkfCV9IxwAgsar9f+rO+YzPgQn2CR8LjozLA9uRsv/t3rr5lG99CswJIwqU9M8DK/W6/jGyqvnII3kJah0jCTO4zwH28b7+iLqW++u/cQnYMScIgcDTALLBvv8Uhob4f3NtCC5lJwuDENMC/h2+/YnOevoi+2kJqIUrC6xs1wMPZbr/LiZu+BaLZQh/BSsL1bDXAg6Zuvx/4mL7yh9hCDUVLwuO9NcDmCm+/eJaWvrdx10KlyUvCWjI2wPjDbr816JK+XEvWQuVNTMLJZDbAexhvv59ykb75MdVC3sxMwoO6NsC/v2+/pf2Ovq4Q1EJxTE3CKgY3wEBJcL+pzoy+1vHSQs7HTcLsPTfAA35xvz9si74909FC2D1OwrqEN8D2PHK/32+JvsS20EIdrk7CKe83wLaBc79Xe4a++pHPQmQhT8JLXzjAnEN1v7Z2g76Oec5CfpNPwojDOMAsWHa/Pp+AvrZZzUIbBFDCYS05wKxDd7/SG3u+OjvMQg5pUMILtjnANuZ3v5Tkcr4aGctC99hQwugfOsCvYni/XoRsvgD6yULVQ1HCcY46wMOUeL+ktWW+JODIQnKoUcIkJjvAfc14v6VYXL6LxcdCtQ1SwpWsO8An83i/VgdUvo+kxkJhcVLCWUg8wOCyeL/BOUq+h47FQnrNUsJWtjzAfuF4v0p0Q76SccRCJiZTwrpHPcDY03i/2mQ6vvZYw0LkhVPCEM49wI6keL+e+zG+80XCQrrrU8IeVT7AJqp4v+ebKb6WJcFCE0BUwt7jPsA/0ni/V84gvnIXwELIl1TCeTE/wE3WeL+6/xu+UwC/QhzsVMJ1yj/A/Qh5v2WTEr5Q4r1C/TtVwrE8QMCJOHm/oowLvs7YvELOl1XCK79AwM31eL/7ZQO+VsG7QjTrVcLgGEHAdvZ4v72x+73Js7pCdjVWwsd+QcCqX3i/ZM7uvTqauUIJeVbC5spBwN0AeL+aO+W9go24QlXFVsKzIkLA4fN2v77w2b2ygrdCixtXwndwQsDRMna/XA3QvYV8tkKqYVfCur1CwFh2db+YQ8a9H2a1QiGmV8JG30LAuTB0v6+nwb0SXrRCfv1XwuMuQ8Ck/nO/yNa3vT1Ws0KTTFjCklJDwGBec7/sP7O9Dk+yQnqgWMKRbkPANLxyv0Wbr73sSLFC2OFYwgycQ8D+qXK/vAiqvTkzsEJ8NlnCt8hDwNxMcr9jd6S9rCCvQl13WcKUxkPAy5Zyv3/SpL2TJK5CsblZwi7JQ8BSK3K/zV2kvdMbrUKTAFrCZ9pDwE41cr/3R6K9ZwasQt48WsK01EPABYxyvwoWo72q+KpCDn1awsDwQ8D4x3K/7b2fvXrxqUL1ylrCTr1DwII3c7+EKaa9j96oQuoSW8JH3EPAzzVzv/tgor0s2KdC+WJbwhO2Q8DxUnO/rBSnvX/GpkJZo1vCL5tDwESBc78ebaq91a6lQqjwW8JxUkPAuQxzv6cms71ajqRCojtcwmVIQ8BUcHO/moS0vfSHo0LukVzC0PBCwKYWc78LFr+94GmiQiTpXMLax0LAaDpzv8YkxL1cUKFCwkFdwtSAQsALenO/hezMvaY+oELxkF3CkUlCwNKDc7+NstO9LC6fQgfzXcKLBkLAmk9zvxrO270MC55CS0lewvHJQcDBGnO/MR/jvYb+nELSiV7CYIZBwBV4c7/Yjeu97eSbQh/2XsI1R0HAiiV0v26d873PvppCYGFfwtsDQcBClHS/kxb8veScmUL9s1/CrtRAwAS3dL8z+QC+i3qYQgYZYMKThEDA08V1v28sBr5zZJdCen1gwrQ7QMDVpXW/i58KvuFFlkJh6mDCX/I/wGgndr+8Rw++UCSVQslBYcK42D/AWJF2v/X6EL4h/5NC7bJhwr+rP8CYPne/5fMTvivxkkJqKWLCilM/wADJd78MkRm+lsiRQmKFYsIuNj/AC2V4vySSG76ep5BC5/diwqQpP8DYo3i/cGwcvrFfj0KTZWPCS/U+wKweeb/k0R++H0COQnjQY8Ks4T7A7+l5v1hKIb7iOATDlH6dQC1QQr/6QTE/eM4sv2DoA8M79pVAfCVEvx8VNj83jTC/45QDw7rAjkA9bEa/w5k6P/6pNL8sSQPD0bmHQEXGR7/i/z0/MW03vz8AA8M5oYBAwXxHv5rVQD9UTDi/DLcCwxXackAttUW/4/VCP5FXN79vbgLDcf9jQBLmQr+fZkQ/sgo1vzkoAsMSJlVAQvI/v2SURT9wfDK/FuABw5D9RUC8TD6/6VRFP/y0ML+wlgHDo9g2QJuJPL85YEQ/EYsuvwJMAcOkOSlAUKY6v1ZjQj/83iu/YAMBw2mHGkAMujm/MRxAP7gTKr/NtwDDrfgKQA/tN7/2/z0/4XsnvyVtAMMOtvc/jZQ4v9j+Oj9+ACe/eSQAwwQ52j92oTi/VcQ4Pxw2Jr9upP/CT7O6P/PFOL/MKDY/Dl4lv1YH/8Kv4pw/+To5v5VfMz+awiS/tlv+wg+Ifj984Dm/Am4xP4+mJL88vv3Cly1BP1/HOb9slC8/btojv0Af/cKq8AM/38U6v0pyLj9DYSS/VnL8wvMjmT5kXTq/AEwuPw/uI7/Uy/vCVM6GPXjfOr9GwS4/G5gkv2QW+8LfACi+D0Q6vw6IMD99ryS/gmb6wj6WyL7cuzm/I8EyP80DJb8As/nCbxgcv/NVOL/PJjU/eo8kv2j5+MIui1O/TDo3vwIzOD/3niS/pEn4wnXGg7/IvTW/Rh87P9E/JL/sjPfCzS6ev4xqNL/McD4/lyokvw7S9sI4ULm/RH0zv0jeQD8hIyS/GiT2wqfr0b/5ozC/0opDP71FIr9+X/XCpNHrvzlZLr9UqEY/9Bghv7Cf9MJb8QHAUQUrv5tKST9rsR6/TuPzwoK/DcDGHyi/+Q9LP0NnHL94LPPCowkZwIYgJb+680w/qwoav5Rw8sKJAyXAqGwiv/eETj8r2xe/cK3xwhrqMMAIcB6/2x1QPxtjFL/g/PDCYk08wOvFGr8HE1I/N1gRv85D8MKj9UbAHDwWvytwUz9uPA2/7onvwpRtUsBImRK/I0dVP6opCr/8yO7CbOhcwICyDb/IMlY/TY4Fvz4N7sLa0WbAU40Iv1XpVz+R7gC/ckftwhTmcsCVUgS/8wVaP9Se+r6+j+zCqvd8wLDU/77DJFs/HnvyvmfT68L5JIPAhHf2vgSgXD9J+em+gxLrwrQLiMBNbO2+CH9dP2J64b4aTurCkeCMwCLi4r5wLF8/lOHXvvaK6cJnspHAnPvavpPKXz/GZNC+INPowl3vlcAyzNG+j6tgP0jDx77wEejCqDibwAP2y75axGA/rxfCvglK58KzsJ/AB6bEvlLuYD8yBbu+npjmwniso8Aphr2+GWNhP1ZBtL5AzOXCshepwNJvuL6uW2E/aEyvvjEI5cJuEq7A44O0vg1SYT+Cequ+kFXkwtw0ssD5KbK+fkhhP5Uvqb7qk+PCehy2wJkIr77xH2A/DcGlvujZ4sKESLrAvhStvupiYD9d9aO+cQziwk4Cv8ARo62+KxxgP5NmpL41SeHC7l3DwIf7rL5Qdl8/T4yjvseU4MLb8cfAcp6uvl1QXz8PE6W+UtnfwiNvzMDR2LC+ErleP0YEp76SF9/CZVDQwMdUs75U914/SX+pvmJU3sJF79TAVe21vmV5Xj9+06u+EKfdwizj2MBWhrm+JJBeP9ZUr74L5NzCYU7dwP+gvr48NF4/sSG0vqMk3MKQpeHAqeLDvi5QXj82Q7m+Km3bwqs85sCwS8m+Hm9eP1mPvr7Tq9rCNCjqwCVa0L6Eyl4/e5LFvlT02cLY5e7AsUPWvvyyXj/kTcu+bSbZwsTu8sD9Idq+RoBfPx9vz74Ub9jCDnz3wLl74L5ueF8/kKfVvqSq18KF3PvARRHmvk+CXz/6Ktu+4PnWwjTb/8DuPuq+0L9fP7Vm374gKNbCOBECwVqG7r59SGA/r+LjvlhX1cLbPgTB6vzxvo3dYD+jmee+dpnUwuOABsEDhPS+Gg5hP6Y06r5F5tPC8XwIwZ9W+L5rNmE/bxfuvgMR08IQ3QrBgo/7vilTYT85XfG+OE7SwpQHDcHSAv2+fuZhP7Ea876TgdHC6asOwS/J/b5sDmI/hPXzvvvF0MJrFhHBvvX+vkCzYj+GdvW+JXnFwvLSvUKZ8U2/jMAgP2RMMb/Qe8TC07C8QqvVUL9evSU/Qz02v7B7w8JojrtCqnFUv5JIKT/jVju/I4PCwv9tukInC1e/f3YrP5LlPr95lsHCX0q5QmrPV7/jtS0/OqxAv1KqwMJAJrhCwHVWv83VLz9dQkC/6ri/wgH9tkIFIVW/oYgxP6iqP79/1L7CRdG1QskMVL+lyzM/05E/v4XpvcK2orRCkIdUvwYCNT9cl0C/1P68wqV0s0JD+VS/avU1P9B2Qb83HLzChEmyQld1Vr+mEDY//gdDvyYwu8LvGLFC0dtZv7IBNj9HgEa/wUi6wrTjr0Jo61u/zvU0P5wmSL/BarnCprquQg6mX7/HWzM/KURLv6qEuMJaka1ChcBjv3XzMT9x3E6/fpm3wmZfrEI7Jme/w/0vP7V0Ub+YsrbC1S6rQoAhar+/ay0/kk5Tv7W9tcLMCapCLDxtvx+wKT98r1S/sNm0wpDaqELzGW+/QDgmPyLhVL/u8LPCrK6nQqlicr92IyI/SjJWv8H/ssLsiqZCshhzv/yPHj+HF1W/UB6ywgZopUKoanW/R+AaP3KOVb+kNrHCnj+kQjvwdr+sixc/xl1Vv3dGsMKoF6NCEIJ5v9TMEz+M/lW/UGWvwmbroULxsXq/qRwRPyfDVb80eK7CAMmgQug2fL+fAA4/xp9VvwGdrcI+rJ9CYT19v77xCz+2ilW/AbyswlCHnkKRsX+/3aAJPwO8Vr9Uy6vC8GSdQn2vgL8Nnwc/rU5Xv3T7qsIWQpxCMKWBvyJlBT8J/Ve/bByqwvgnm0L1l4K/SEYDPyyxWL+lOKnClgCaQjA3g78+VgI/02ZZv31TqMIU4phCCf+Dvx1DAD+Yx1m/7oinwg3Il0Jqu4S/QfL9PoSBWr8HqKbCRKGWQhuahb+ipPs+bpJbvyLZpcJ6epVC1q6Gv3iX+T6GIF2/9AWlwv9elELHs4a/Guj5PsVBXb/VNaTCIEeTQqwYh7+F4Pg+7L1dv1Bno8I/IpJC/ZmHv2Z4+T4l7F6/A5aiwhQNkUK1p4e/6675PoMXX78xuaHC8+6PQp+Ch78Nwvo+6x1fv37voMKj1Y5CNHSHv/5s+z4RM1+/iBCgwnrCjUIceIe/c6L8Pl+VX7+pP5/Cy7WMQs5Ih79e2/w+N0dfvxRxnsKYm4tCrFeHvw0e/T56eF+/Xpadwj2JikIWf4a/wsP+PnNBXr8juZzCWHuJQpg9hr+Jc/4+SKddv9H/m8LmaohCiNKFvzTt/j6n9Fy/FCWbwjlWh0KCcIW/yPr+PiM1XL+xQZrC1FCGQvzkhL/elP8+k0tbv1GMmcLjVoVC+yKEv9g+AD96DFq/jreYwqNNhEJXAoS/TZf/PnCKWb+kwpfC5TuDQrsdg7/M2v8+ddlXv3cUl8KvQoJCyYCCvxGV/z6ckFa/JEqWwm9HgULLBoK/Fa3/Pn2nVb8WgJXCIlOAQo5Pgb8klgA/IqpUv2+slMJHgn5C18+Avx6xAD+4vlO/meCTwl2ffEKKTIC/go0AP2mqUr+cHJPCRI16QtMFgL/v4AA/+U1Sv9BRksJWmXhCbrJ+v4vAAD/h61C/rIiRwpK6dkI1DX6/uQwBP3Z0UL/trZDCU8d0Qj2cfb9l3wA/K+5Pv1wAkMKF73JCO3F9v85CAD/gb0+/hk+PwnsMcULoiH2/jBUAP2JuT78mfI7ClBNvQrQffr969/4+TK1Pv067jcJMJm1C1459v7zU/z7PXE+/ydaMwm5oa0JLGX6/TFsAPyQgUL+XHYzC8GlpQlkSfr8RLAA/1P9Pv4dSi8IwimdCAot9v5+oAD/av0+/8IyKwg3CZULmjX2/JvwAP8bvT7/lyInCxNpjQovCfb/pzQA/6wlQv9QSicLtEmJCq2N9vyp1Aj89klC/dTmIwoo8YEIbbny/2rUCP+XFT78VYIfCRIxeQqeYe7/giQQ/zPBPvyiphsIqx1xC5cJ6vyXqBT8Y3E+/es+Fwmv5WkITw3m/DAsIP0oDUL9E/ITCVkdZQvcbeb+wSQk/9wdQv5IZhMIClFdCTIx3v2DOCz8y0k+/NVqDwtPNVULsEHa/Iw8OP1aJT79Jb4LCl/xTQpTIc7+B2hA/zrlOv+HH0sLq07VCMpIpwOEoa79bq/a+PWHTwq5dtULB+SXAGgRkv7sZCL+gANTCqOq0QkZCI8D54V2/U2MRvwGm1MLAdbRCY6IhwCaLWb/9qBa/UE/VwhAEtELjaSHAbrhVvw5WFr/1/NXCMI2zQnjhIcBAMFO//aUTv7ef1sJlE7NC/IciwKvlUL95VBC/70zXwiaYskKoIiPAFLpOv9lDDb/C/dfCNRyyQnimI8B3hk2/AN8Kv5uj2MIBorFCVCUkwN1MTL+2jwi/AkrZwtgwsUL2biTAl9tLv3xOB79k6tnCMbSwQvhtJMBMZ0u/IjAHv9eU2sL+MrBC/Y0kwEZdS79dsAa/DkPbwvK7r0JrSyTAIXZLv0y7B79k5NvCk0CvQv6gI8BeZUq/WgMKv5x/3MKzwa5CYjIjwDZ9Sb8NbQu/qCzdwvZDrkJLnyLAySdIv6lCDb/uu93C4M+tQgYgIsD4LEa/NZUOv7Na3sLJS61CNq4hwJw4RL9Wsg+/6vbewkTNrEK2ByHALy9Cv2iSEb9qkd/CXFGsQjjgIMAgbkC/F5oRvyYw4MKQ2qtCBUUgwCaNPr9fVRO/jL/gwi9bq0JK1x/AkBg9vwSCFL9cWuHCjtuqQvI4H8DU+ju/CoYWvzf64cJ+WKpCUbcewGbKOr92Exi/tY7iwlTfqUKTRB7AL5k6v6e+Gb9DKePChGWpQvq5HcC1ETq/0agbvyfC48Iv5qhCvyQdwGjeOb+32h2/LUzkwvJlqELyfBzALVY5v6k3IL8T8uTCw+SnQkzuG8By4Dm/F5giv1CE5cKAZ6dCRl4bwGr2Ob+01SS/9RDmws3kpkIvuhrAvpI5v482J7/moubCbGGmQgEpGsC0/zm//p0pv/5L58Jk5qVCbJMZwH5NOr/lDSy/CNXnwtldpUKt9RjAnjQ6v9R4Lr/BaujCZdakQgpaGMC5/Dq/cTcxv44E6cL8UaRCTNgXwIanOr+pHjO/OKzpwibPo0JsfhfAuno6v4F2NL9CSOrCykejQiLqFsBO/Dq/BQM3v77W6sLzwqJC4mUWwDUVO78kJjm/CnLrwjIxokJFNxbA9/87v0ZGOr/aAezCa7KhQpwQFsCBbTy/KhI7v2SL7MIMLaFCh7oVwGMKPb/rtDy/tCbtwkOpoEL+hxXAR949v5fePb9Yxe3CwhqgQpVVFcAhPD+/xkM/v6hO7sLpjp9CsGcVwAPuP7/qRD+/yO7uwloHn0J4dhXAzsNAv/NiP78og+/CyneeQq1xFcCMpEG/s9Y/v+Qa8MJD8Z1CNHwVwE0iQr+X4D+/6qTwwk5qnUJvgxXAWXhCvzbnP796QfHCRumcQuqOFcCFjkK/68A/vx7K8cLcYZxCf3QVwGB0Qr+lI0C/zGLywvnKm0KVsRXAXzBCv+cIP79s6/LCRVKbQpC9FcCA+UG/678+v4h/88IOzppC6NIVwBd/QL9Nxz2/4BX0wo9LmkLp3BXADXM/v6gsPb+on/TCf7OZQsGzFcAyRT6/RFY9v4Is9cJoM5lC5sQVwHQJPb/+iTy/zsf1wg+mmEK+uRXAaPc6v/nWO79YXPbCDiOYQjuuFcBkhTm/7mg7v7zp9sKem5dC8L8VwCQBOL9PfDq/Tmn3wlkOl0JWvBXAvU42v8HSOb/IEPjCW5CWQqG3FcBnvzS/Zjw5v0aX+MJ7FJZC/8YVwJ7vMr8YOji/ziP5whyElUK50RXA1VAxv5ZfN7+QrPnCHgKVQqCuFcCCJC+/nP82v+wU+sKphJRCgn8VwF0lLL8ZdDa/6KX6wnr8k0K+rBXAn+Apv+bKNL8aGfvCiHKTQja2FcCXPie/zYczv4qd+8JZ9JJC84QVwAw4JL80/jK/dDD8wq97kkLNtxXAgDEivy9dMb9mrfzCPv6RQn53FcD4oB6/JdAwv6IS/cLSgJFC2bUVwLQdHL/D0C6/rJT9wsENkUKHixXAzkkYv3XQLb9aC/7CS5CQQo6eFcDBYBW/HUwsv1xv/sJ+DJBCYGoVwLevEr8y6Su/buj+wrehj0K9exXA5QIQv7WEKr8gO//C1CSPQtErFcCjuwy//EMqv0Sz/8Ioq45CsDkVwCRqCr+qEim/+QAAw5s0jkKtTRXAZQ8Iv9THJ79YeM/C56nCQvHOoL+tYJA+zFxuv1rLzsLOX8FCbWKhvzQmnz7Y4XS/iyTOwlsWwEIWB6K/KiGqPks0er+8g83CF8++Qmcqor8mQbE+fhp9vx7vzMI8hL1CIwShv2BAuD5ERn2/ZlrMwuA5vELG8Z6/FDe+PrMfe7/zwcvCNuq6QgYQnb+EKsM+b/x4vzc4y8KCmLlCfTKbvxoLyT5sMHe/uqfKwlBFuEJaRZq/okTMPp5kdr++F8rCBfO2QlBtmb/tBc8+P5l1v12NycLVp7VCvniZv97Ezz4J8nW/l/nIwrtTtELKcJq/ktXPPq34d78Ub8jCoPuyQhLNmr/ABM4+lRd4vx7sx8LPsbFCZiicvzwEyz5e3Xm/wl3HwrJlsEJus52/dqfIPu4/fL/Gx8bCsRGvQlc4n7+X/8Q+hiB+vxo+xsKswK1CsW6gvwJlwD6vAn+/ZJ7Fwtp7rEI0uaG/ZCK6PvNvf7+eEsXCzCqrQseRor8GCLQ+3vZ+v7Z+xMJ63qlCuRukvxxArT53qH+//+HDwvCbqEJqVKS/efamPgnFfb93VsPCelmnQoyQpb/+yaA+NAB+v8LCwsK7EaZC51+mv9TTmj4SaX2/SSvCwtLLpEIb8ae/peSTPhX6fb/in8HCfH+jQnWkqL/gP48+3J59vzEIwcKzQKJCirSpv0pbiT6YgH2/ToHAwnUHoULxlqq/DC6FPqKrfb9g87/C/8SfQlcArL9W5IA+Vtt+v0VRv8KKhZ5Cu+Wsv3mRez7Acn+/Ude+wiRInUK8KK6/kkFzPrstgL9LRr7C/BCcQiJHr7+V8Ws++5SAvxqyvcIC0ZpCCvqvv3ddaT4TCYG/9Rq9wl+TmUIS97C/x85iPlNggb+SnbzC91+YQqy8sb+Ut14+LL+Bv7ILvMIDH5dC1Z+yv9/yWz61X4K/2Yi7wkbflUIDvLO/NrZYPh8vg78o/rrC2KiUQnqys78kolo+NlmDv3F6usLZdZNCLAy0v34FWj4OpoO/hfK5wtI4kkIuY7S/QkRcPvM9hL9narnCUwuRQt9ftL/un10+b1+Ev7nTuMIS0o9Chgu0v/KmYD5MWYS/BUi4wpShjkKIw7O/QpFjPtdchL+0r7fCRXWNQsOis7927WY+YJWEv6Ubt8KRT4xCFVKzv4P0Zz5vXIS/7JK2wmAci0IrWrO/5NNoPnd8hL/D9bXCEe+JQmQxsr+sDW0+/baDvxpdtcLpyIhCAp2xv952bT4PJ4O/wOG0wjSeh0Kc/bC/oRxvPuysgr9WR7TCWXGGQtNasL+znW8+YhGCvwuls8KZVYVCw4Gvvz7YcT6Wa4G/xCmzwmY+hEKfjq6/vJhzPgefgL8DkbLCKx2DQokkrr/ranM+nC2Av2TkscKu8IFCXvKsv07MdD5gMX6/+myxwi7ggEKSLay/1nV1PqjDfL9r3rDChJZ/Qgc9q7+ORHc+Jzl7v1pPsMIwf31C2xKqv/YdfD7e1Hm/K7Wvwsole0JgWKm///J9Pha6eL8dI6/CHBJ5QsZ6qL9WMn8+yzx3v8acrsKrzHZC3Oanvx4KgT4KpHa/4AmuwgyhdEJeBqe/arOBPmQmdb+We63CQolyQgtSpr+7JoM+Pk50vy7XrMKaaXBCnuqlv8rTgz4cw3O/xWSswjxdbkJRoqW/tkeDPkT/cr9r6avCZEtsQjGHpb/QhoM+kuFyv/9Lq8KiIGpCzIKlv8sRgz68rHK/EsWqwhIGaEK1FaW/CB+EPjI7cr/nE6rCDBFmQl0Vpb8CgIU+Qr9yvw2PqcKT6GNCT6OkvyvmhT6xBHK/e/mowoXVYUL5KqS/iEiHPjyccb+/aqjCANxfQinjo78ut4g+gpdxv37dp8KkxV1CB9ejv55AiD55U3G/TlinwpfMW0KzQqO/89qLPsuFcb8HtabCEMBZQjyOor9wq4w+Zm9wv3IMpsI861dCwqqhv4J6kD7BFXC/tIGlwi/6VUIlvqC/TVeTPkFPb78Y4KTCjvdTQlHen7/Ui5c+Jhxvv91BpMLkHVJCHTifv5ncmT7uqW6/x4mjwsZAUEKW8J2/dEqfPk4Ubr/a/aLCRk5OQt6xnL/kmqM+AyZtv0ZBosKLWExCoPyav1X/qD5qqmu/d7gCw0h9dkIrAEu+AMtmPxqkQb6YTwLDAMZ1Qiytfb4iqGg/TFtzvr/iAcM+EnVCCr2Vvqeraj84iJC+X30Bw75jdEJxVKO+jbNsPyWFnr6MGgHDB7hzQh4nqL7bU28/xSekvnu3AMMMCXNCutCmvs5mcj8y1qO+H1IAw25PckJjGaG+noZ1P1Ien76a3v/CKZNxQgrsm77mNnk/SA6bvrQY/8LmzXBCia2XvpPjez/zkZe+ukz+wrQJcEK63JO+NxF+PyFWlL6ihf3CEElvQgvYj74E2X4/wniQvoK5/MIShW5C3wuPvlptfz+20o++Lur7wm6zbUJ1wYy+JBF/Pwxkjb5uIvvCuutsQlwkj765Cn4/3IePvvha+sKqJmxCXduRvnE2fT9KDZK+Tob5wj1Qa0J2hJS+wqN7Pz1LlL6gtfjCDn1qQsysmL7/2Hk/2PeXvgTV98IRuGlCEvWdvuyEdz+4lZy+Rgf3wkncaEKDLqG+nHV1PyMun76yOvbCcgNoQux3qb4F9XI/A62mvlZd9cIXOGdC7nutvnUfcT/QEqq+qo70wp9qZkIO47S+4b1uP3ehsL6kr/PC94plQt1cvL5a+Ww/1G+3vmTN8sInrGRCQsDEvu/Taj+K+L6+BPfxwlTPY0KjG8y+NahoP4Fsxb6WD/HCVvpiQn491L59mmY/yqbMvuA98MLwJ2JCewXbvp5BZT88ztK+oFzvwhJIYUIgwuO+qONjP/Hf2r56e+7CWmlgQhMI7b4O22E/ziXjvkSs7cKujl9CsXvzvo43YD+oxei++sXswj2tXkLrMvu+yuZePxjK775l7uvCx8RdQrTg/74jH14/Igz0viEL68Jp9FxCZb8Cv7uXXD8w2Pi+I0Pqwt0dXEJ9CQW/N9ZbPycA/b4oX+nCry9bQmNnB78DQVs/UrMAv+SF6MIaPFpCgVkJv5r7Wj8HkQK/KrfnwlViWULuBQq/w39bP6dhA7/i4ubC2YlYQnaNCr9hJlw/ZBcEv34X5sJil1dCeU8Lv3QwXT8GJAW/Dz3lwq+7VkLEsgq/Oe5dP1u7BL9ma+TCFt5VQiu1Cb+SgV8/sSsEv3Kd48It71RCapgIvxZCYT9/hwO/fMjiwsYUVEL/qge/aIxiP2vxAr9kBOLC+jxTQvorBr978GM/qs0Bv4ow4cL6UlJCOo4EvxxEZT/ChAC/GF7gwmZ+UUKp7QG/b+tmP56S/L50id/COp9QQuALAL+s52c/ZkL5vpjG3sIezU9CJ976vhL8aD+KgPS+QPzdwgToTkIOoPe+RghqPyu78b5KGN3CJBNOQiOc8r5oPGo/Ur7svoxr3MItZU1CDZ3tvlaKaz+ITei+Z5nbwuR/TEJ+2+q+lrtrP8uZ5b6pttrCOJxLQvAF5r4yW2w/wf7gvtAF2sJO4kpCMpPivhNwbD+ki92+k0LZwv4eSkJSjd++FVdsPxlz2r74idjCW15JQhe/3L6Hn2w/nr7Xvpiv18LPgUhCr6favtOlbD+/pdW+ounWwgW/R0IY2di+CsNrP2Fy0761ONbCQ+BGQmwI2L4iUms/XnDSvnKB1cIYJ0ZChdvVvhqBaj8Z6c++PsPUwhB0RULmRdW+sclpP9EFz76f9tPCjqlEQjr70758TGg/UxvNvn9M08L/9UNCR8fTvsdMZz9MfMy+HZzSwg0zQ0KlkNS+orhlPy6azL7z4NHCpHNCQmDv1b4KSGQ/eFnNviAq0cJRokFCTuzWvugoYz8u2c2+rWTQwlAFQUKx39m+hSRiP1FS0L5grs/ChiNAQqaa277H4GA/DHrRvpbxzsKSZT9CFinbvmuCXz8JctC+LjrOwsO5PkKHN92+2B5eP4Xa0b5jhM3CmP49QrUD377T3Fw/mA7TvpjdzMLCUz1CGOjgvrF8XD8bvdS+whXMwpyvPEKC1+C+2kJbP9Mg1L42WMvCugM8QuwW4r7ySls/zVvVviuqysIRPTtCXC/jvoMfWz+4Wda+burJwruUOkIiReS+1gRbP5Nc175AKsnCXfg5Qhdr5L7WB1s/9oLXvtRgyMJ5TjlCChrlvnv6Wz8UnNi+DbHHwnarOEJYKOS+6GZcP/Dg17537cbCIvM3QtLO4r5AJV4/DljXvs7eA8OOYHdCMqwQwJTFJL+ZXka/9woEw7codkLywAzA4kYdv3aIUr/6NwTDd/x0Qt/GCcCbZRa/yxNbv9FoBMNM1XNCNgwIwD1cEb9daV+/y5wEw8CzckIZnAfAmaIMv/SPXr9u0QTDBpFxQi7wB8Bk+wi/PjBbv3oCBcOUZXBCm3IIwFHcBb/fale/qTYFw8w3b0KA3QjAsMECv/cOVL8kawXDmgluQtZECcDBNQG/GKJRv3+cBcNo2mxCj2UCwFra/77a7E+/+8gFw8bMa0I+AgLAeob/vvHITr9C9QXDwZpqQhznAcDg6/6+o+NOvwslBsNiXWlCCJgBwJ12/r52JU6/nlUGw1I2aEJzwAHAiP79vgIXT7/8fwbDPQdnQkl2AcAuzPq+KJpQvyKlBsNazWVCnmYBwPQk+L5Nf1K/NNIGw2SaZEIy/ADAMnP0vnbpU78D8gbDznNjQlo9AMA/DO++QG1Vv4AaB8N0L2JC3+v+v/NL6r71QVa/qD8Hw5H3YELQnP2/ZIjkvhtGWL/UYAfDeMhfQlYw+79C4d6+lhhYvxSGB8MknF5Chhb6vzJ/2b6cFVq/SqIHw95iXULU5fi/FZPUvpWFW78NxAfDXClcQk1j+L/qXNC+6ahdv1fnB8MC7lpCfWv3v3z5y76zDF+/IAIIw0TDWULGV/e/+APJvikMYb9FIgjDY5pYQncG978gesW+ZAFjv5U9CMMNa1dC8gr3vxQ7wr7EYWW/9VAIwwAxVkJ0Dfe/UIy+vigPaL9ZcwjDZgJVQkrN97/X6ry+hrNqv4KJCMPW0VNCCg34v181ur6oKW2/PJ4IwzKjUkJC3/e/d4G2vuF/b7+vsgjDUmtRQnWE+L/+jrS+Xzhyv+bSCMMoSFBCsQT5v0rZsr6yfnS/2+QIw5IPT0Itbfm/pmqwvisgd7929wjDOdRNQkNB+r+IYq++rp55v+QOCcNFnkxCoVH6v170rL6/jnu/VCsJw6lxS0LoZvq/WvGqvkU6fb92QgnDRTtKQhXg+r8YaKm+2mJ/v1BRCcMcBUlCwU/7vyUzqL63nYC/hmYJwwzAR0JOZPu/LTinvgcSgb/6cgnDnZJGQvMZ+78olKW+2F2Bv8GDCcPMXkVCBpz7v+0Bpb5NI4K/BZcJw38rRELFvPu/bp+kvmFsgr9UrAnD5O1CQtaG/L9wp6W+yOiCv3q3CcOes0FCnMn7vyrgpL4/YoK/DMsJwyZ7QEL0rfu/3emkvhdAgr942wnDLjw/QodT+786t6S+iO+Bv3/sCcPQBj5CV/j6vz6dpL49lYG/XfUJwzjWPEJdQ/q/mJqjvuAvgb+TDArDSK07QkiQ+b86qqK+TcaAv2wWCsN7bjpC2ZT4v7+9oL6da4C/tCkKwyElOUI9Mfe/+8mevrxIf7/mMgrDvw44QpA69r9zFJ2+hnd+v1lFCsMg5DZCnHD0vxDNmb7gB32/IVUKw4zGNUJnx/K/ZLCWvtG/e78RWwrDyHI0QsxK8b8ubpK+Xp57vwhlCsMnXTNCNYDvv2BXj77YC3q/BnsKwwgcMkKnqe2/eI+KvpiEeb9QjQrDDf8wQnzo67/wBIe+WlF4v/GTCsOrxi9CqeHpv4lOgr7aVHe/35kKw5mdLkL0cei/4X59vjPDdr82sgrDEYYtQg975r+bI3S+nth1v9u4CsMAZixC1tzkv7YPbL58MXW/s8QKw181K0LFFeO/rnljvsledL920ArDRwkqQsa24b8nYVu+/St0v87LCsPK9yhCO9Xfv/X7T7549HO/hNgKw7fDJ0L8qd2/hotGvh6Scr+q2wrDGJsmQuh3279kZDy+YVRxv3vfCsPZayVCRq7Zv26PML4RUnG/WewKwwlgJEKJPNi/OaIqvrlCcL8l+ArDSj0jQupi1r8eRR++3fFvvzPyCsNNDCJC+YPUv/g+Fr4X6m6/HfQKw/4JIULzqdK/VgQMvm06br8N+ArDlu4fQnR50L8gJAO+yIdsv3/1CsOErB5CtKbPv8Wu973i7my/LfsKw2itHULFN86/WJnrvVTda7+28ArDDIMcQvv5zL9+mde9ESpsvwL3CsNwaxtCONvLv7RLzr1nS2u//ekKw59HGkImV8q/sz7AvUpMar8i9NnCC3fIQhNgH0D/z1O/JfsdP06a2sKt9shC8CEgQBhXV79rEhw/YC7bwoN6yUIVsiBAf7RZv8uLGj/6z9vCs/rJQuwaIUAEmlq/BCkZP5563MLVfcpCXOMgQEqiWb81vBk/CCndwgj6ykLUdCBANPNXv0PzGj9K1t3CT3XLQrbGH0A9i1W/1O0cP32E3sLb78tCVRYfQLbKUr+7zh4/1TvfwuNtzELnfh5AgzhQv4dUID8P49/CpevMQv6oHUDc002/w+IiPxub4MKuVM1CjPEcQOc4TL/rOCU/NUPhwsjVzULDNRxAjuBKv9e2Jz9L7+HCFVLOQn6BG0CAPkq/rVgqP4uW4sKm085CsR4bQOoKSr9c2Cs/BkvjwmVOz0L18xpAi15Jv/hFLD8t5ePCdMzPQmSDGkC0V0m/QA8uP5aX5MJZQ9BC4VUaQLCISb893C4/izDlwgjD0EJ9PRpAPFRJvxgsLz8kyOXCwzzRQoQSGkDm8ki/0bYvP7d05sKcutFCoQkaQHyHSL/jsS8/eQjnwhc10kKw7RlAxKhIvy8xMD9PqufCiLrSQm4LGkCIpki/d7YvP2FQ6MJmMtNC5UgaQO5CSL/JlC4/COrowiKx00KRixpAJxdIv+lzLT/ah+nC3C7UQsUaG0A3aUi/m0srP0ci6sK5rdRCVKgbQBJvSb+rbSk/wMDqwk0u1UK7MhxA099Jv05lJz/ka+vCU6DVQp7ZHEAD5kq/ISAlP8wJ7MJBH9ZCv5IdQCaYS7/2ciI/wLjswsab1kJfLB5A4gxNv0CJID8CVO3COBfXQgTFHkCKt02/SlweP+z47cKOitdCdDsfQCg9Tr+urBw/ipLuwtIN2EIvpB9AB8lPv0GMGz8AQ+/ChIvYQmMDIEDh11C/2WYaP7ze78IC+thCW1cgQBHcUb86ahk/CJbwwl932UIImCBAtWlTv5DnGD8kPfHCnuvZQg7KIECHRVS//WQYP0Du8cIAaNpCfZkgQBugVb/umhk//J7ywu/Z2kJ8ySBAC4xWv+AkGT9MRPPCplbbQjirIEAj+le/excaP57y88L2w9tC53UgQN9QWb+pYRs/3pv0wlQ+3EK2KiBAcoZav1D8HD9gLPXC7KncQrQWIEAEK1u/L4UdP0r29cLpH91C3g0gQO65XL/zLR4//pn2wnmU3UKv8B9ADwdev4EVHz/gMvfC/gzeQqWUH0AQ1V6/29ggP/7y98Izet5CtWEfQHB3X7/a4yE/UpP4woXi3kKsUR9AgHJgv657Ij9WP/nCd1jfQg4mH0Cz52C/BVojP7re+cIizN9CcggfQI5SYb+J+iM/AIX6wuJI4ELa6B5AhgBhvyJjJD9iKfvCj7zgQqrSHkD+kGG/wPEkPz7I+8J0IuFCEq8eQNuiYb+RjSU/wHX8wsKa4UI8nR5Ayn1hv9zLJT+WB/3CXQ/iQqGbHkA6M2C/yGAlP2zE/cIng+JCjkweQM+tX7+VfiY/ynT+wrLq4kJfjh5Ad2RfvyBRJT8AFP/CtlfjQnyfHkDnFV+/X+4kP4jS/8I3zONCvsMeQCG3Xr8jNiQ/RzcAww485EJxwR5ATgtev8wEJD9+lADDPLPkQrXrHkC7WF6/424jP2TcAMMfJOVCju4eQKoZXr93TSM/njwBw3KM5UK5Zh9ALTxfv+u6IT90mQHD+gjmQn56H0A++F+/1achP77wAcO+YeZCeiAgQEdaYL95FB8/5EgCw7bU5kIaUCBAm5hgv8hiHj96nwLDHD3nQgTBIEAv/GC/vq0cP+buAsOYp+dCQSohQB7cYb/xQBs/fD4Dw0sR6EJlZSFAGE9ivyRwGj+JoQPDsX/oQhEiIkDSy2K/6ogXP5T5A8Oh8ehCf2IiQON6Y7+ytBY/aU0EwzJw6UIQriJAqaViv9E5FT/tnQTDrOHpQpkWI0Bic2O/SckTP5cABcMYROpClmQjQJoAY7+nZBI/XUsFw5Cg6kKN3iNAaIViv9xJED8gnwXDyBnrQiBDJECxLWK/R5MOP0zsBcPegOtCWJEkQAPbYb/bOg0/ZT0Gw9/n60IJHyVArTJgv/l9Cj/EjQbDAVPsQnNZJUDE8V+/IX4JP83XBsNYtOxCE7klQIlbXr9hhwc/BCsUw2XJgEGlshtAmvk8v4qnJD8gZRTDyTKCQR1CHUCV+0K/mKcgPzSUFMNBmYNBtZkeQJMvRr/dchw/8scUwwL5hEG9YR9AinNHv+nHGT9P/hTDvmCGQblvH0CIwUa/d1QZP3Y0FcM7wYdBtO4eQB80Rb/myxo/N2kVwxwTiUHVJR5AAbVCvxMJHT/anRXDGW2KQSr0HEDVWT+/po8gP4HUFcPCxItB4NAbQHnuO7+IzSM/8gUWw3UfjUGzVxpAuk44v6s/KD+wNxbDFIaOQeLUGEB1HjW/iPgsP8FlFsOr6o9BGXIXQEOGMb85ADE/rZQWw1E2kUGQ/BVApgYuv5VRNT8HwBbD552SQdXbFECBqyq/IFk4P7bzFsMh8ZNBIeATQHNQJ78WxTo/tBoXwz5PlUFwzRJAHiYkv8GYPT8USBfDrpeWQTLNEUBvBCG/Zh9AP9hvF8Mu7JdBGv4QQAxoHb+5okE/gJUXw5pGmUFuPRBAngcav4cAQz/ZwBfDgZSaQdOSD0Cevxa/cQtEP/7lF8NX55tBuQ0PQBgFFL+CwUQ/iAsYw+FFnUE8nw5A+nERv+4tRT+5NBjDm4OeQTp9DkCSDA+/UoREP3lWGMMazJ9BXEwOQDlIDb8jYkQ/Z3wYw4QjoUHdaA5AmWoMv8uGQz9PmRjDtnGiQTy2DkAXzgu/fRBCPwy9GMMUxaNBjuIOQPVPC7/5KEE/BugYw7P3pEHPJw9A3KULvwtKQD+QBBnDEEymQQWtD0Dwjwy/or4+P5krGcPhqadBHB4QQN4aDr8xzD0/4UcZw2zpqEEadRBA+hoPv9P5PD9ubhnDXimqQQ7BEEBMtRC/WZk8P1eNGcOllKtBCwkRQKbxEr8zkjw/qbEZw/PWrEEDUxFAtW0Vv46ePD/H0RnDhAeuQWmPEUC//Re/yuU8P1b0GcMyX69BWMMRQMrIGr+YZj0/sBkawwWUsEFd1hFAw8ocvzcJPj8UQBrDJtexQQrYEUCVoh+/slE/P8RhGsO/KLNBVcQRQPf2Ib82skA/7oAaw7t2tEFnthFAjRckv7zkQT+CpRrDoLu1QSmTEUCGGia/eV9DP4/BGsN89bZBJB0RQDcYJ79cr0U/fN4awzASuEGsGBFAFlAovyJTRj9oBxvDYT65QdznEED5oim/brdHP4glG8MPm7pBUpEQQJ1mKr9uc0k/eEUbw+vlu0FpNRBAFHAqv/DuSj/eaRvD1wG9Qd//D0AWfSq/FtBLPwKGG8NnPb5By40PQLlcKr8ulE0/AKQbw8F6v0FOSA9AoNopv6pyTj/2vxvDgM/AQWHoDkBttii/rW5PP2DhG8PGH8JBK50OQLKbJ782GVA/Xvcbw/V3w0Eb+g1ADNslv47ZUT+SFhzDt7bEQR+ODUC7WiS/fdVSPyAvHMPcAcZBJycNQGRZIr/ieVM/+kocww1jx0G+xAxA0yYgv8/vUz8lbBzDdtXIQbkEDECUeh2/26JVP7KHHMPjFspBF6sLQO3zGr/YwVU/Kpgcw/xly0E2CgtAwtkXv7+xVj9fwxzD6erMQemqCkDJeRW/2fNWP9fcHMNybc5Btf0JQCBIEr+t/Vc/TPMcwxvFz0ERcglADIAPvxyyWD+gCB3DyEjRQTrQCED5kQy/86VZPxorHcPym9JBBV4IQP/ECb995lk/M0Idw9IL1EHGvgdArNIGv+fBWj8EZR3DiHzVQad0B0Dp3gO/O0JaP8CAHcPw3tZBfQIHQJXuAL/OX1o/pJ0dwy562EGhSARAihL8vpAhWj+Ltx3Dj83ZQVg3A0BjWfe+wfBZPyLKHcOqbttBoAICQFiT8b6pCVo/GPEdw9er3EF5xQBAzZ3tvseFWD/yEB7DF27eQY63/z8meOm+0F1YPxUyHsPKG+BBPmn9PwrA5L7OvFc/7kAew0CO4UFF/vo/wJ/gvhNwVj99Zh7DTvPiQaJx+T/82t2+DrJVP3N5HsN7WeRBbWv3Px3C2r4KWFQ/upMew37z5UHP2/U/s+nYvvjqUj+KsB7DVETnQTGU9D9uYNe+fMZRP0TVHsMVq+hBXujxPxBX1L7ETE8/iuwew/wx6kGUKfE/r5bTvu2MTj/u+x7D6KHrQbfx7j8mGdG+doNMP1BLAQItAy0AAAAAAAAAIQBvdRbKgCYCAIAmAgAUAAAAAAAAAAAAAACAAQAAAABwcmVkaWN0ZWRfYWdlbnRzLm5weVBLBQYAAAAAAQABAEIAAADGJgIAAAA=" + } +} diff --git a/code/tt_diffusion_planner/samples/straight_road.npz b/code/tt_diffusion_planner/samples/straight_road.npz new file mode 100644 index 0000000000000000000000000000000000000000..e4373a6e27f869b8969b0f4948ed7aaa0b6d4700 --- /dev/null +++ b/code/tt_diffusion_planner/samples/straight_road.npz @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:b9979b9007df6ca1190b4b9ef61e24534f862c41b2635d0219dd22c2dfd9d381 +size 10832 diff --git a/code/tt_diffusion_planner/samples/straight_road.reference.json b/code/tt_diffusion_planner/samples/straight_road.reference.json new file mode 100644 index 0000000000000000000000000000000000000000..b2fd89233c6b93ec4f72ee868d05f876fe42e518 --- /dev/null +++ b/code/tt_diffusion_planner/samples/straight_road.reference.json @@ -0,0 +1,887 @@ +{ + "model": "diffusion-planner-p150", + "frame_id": "base_link", + "meta": { + "predicted_agent_columns": [ + "x", + "y", + "yaw", + "cos", + "sin" + ], + "predicted_agent_rows": [ + 0, + 1, + 2, + 3, + 4, + 5, + 6, + 7, + 8, + 9, + 10, + 11 + ], + "force_stop": false, + "time_from_start_s": [ + 0.1, + 0.2, + 0.3, + 0.4, + 0.5, + 0.6, + 0.7, + 0.8, + 0.9, + 1.0, + 1.1, + 1.2, + 1.3, + 1.4, + 1.5, + 1.6, + 1.7, + 1.8, + 1.9, + 2.0, + 2.1, + 2.2, + 2.3, + 2.4, + 2.5, + 2.6, + 2.7, + 2.8, + 2.9, + 3.0, + 3.1, + 3.2, + 3.3, + 3.4, + 3.5, + 3.6, + 3.7, + 3.8, + 3.9, + 4.0, + 4.1, + 4.2, + 4.3, + 4.4, + 4.5, + 4.6, + 4.7, + 4.8, + 4.9, + 5.0, + 5.1, + 5.2, + 5.3, + 5.4, + 5.5, + 5.6, + 5.7, + 5.8, + 5.9, + 6.0, + 6.1, + 6.2, + 6.3, + 6.4, + 6.5, + 6.6, + 6.7, + 6.8, + 6.9, + 7.0, + 7.1, + 7.2, + 7.3, + 7.4, + 7.5, + 7.6, + 7.7, + 7.8, + 7.9, + 8.0 + ], + "valid_counts": { + "ego": 1, + "neighbor": 12, + "static": 0, + "lane": 33, + "route": 6, + "polygon": 1, + "line_string": 9, + "goal": 1, + "ego_shape": 1, + "turn": 1 + }, + "reference": "fp32 CPU (torch)" + }, + "timing_ms": {}, + "num_poses": 80, + "columns": [ + "x", + "y", + "yaw", + "cos", + "sin", + "velocity", + "acceleration" + ], + "trajectory": [ + [ + 0.8299, + 0.0007, + -0.0007, + 1.0033, + -0.0007, + 8.0591, + -0.3468 + ], + [ + 1.6377, + 0.0009, + -0.0006, + 1.0026, + -0.0006, + 8.0244, + -0.0349 + ], + [ + 2.4372, + 0.0009, + -0.0012, + 1.0021, + -0.0012, + 8.0209, + 0.0592 + ], + [ + 3.2377, + 0.0016, + -0.0016, + 1.0021, + -0.0016, + 8.0268, + -0.0312 + ], + [ + 4.0421, + 0.0016, + -0.0023, + 1.0018, + -0.0023, + 8.0237, + -0.0017 + ], + [ + 4.8417, + 0.0019, + -0.0021, + 1.0017, + -0.0021, + 8.0236, + 0.0849 + ], + [ + 5.643, + 0.0022, + -0.0018, + 1.0014, + -0.0018, + 8.032, + 0.0807 + ], + [ + 6.4473, + 0.0021, + -0.0013, + 1.0011, + -0.0013, + 8.0401, + -0.0193 + ], + [ + 7.2494, + 0.0013, + -0.0014, + 1.0009, + -0.0014, + 8.0382, + 0.0377 + ], + [ + 8.0545, + 0.0012, + -0.0013, + 1.0007, + -0.0013, + 8.0419, + -0.0051 + ], + [ + 8.8587, + -0.0027, + -0.0003, + 1.0005, + -0.0003, + 8.0414, + 0.0536 + ], + [ + 9.6567, + -0.0027, + -0.0001, + 1.0003, + -0.0001, + 8.0468, + 0.1724 + ], + [ + 10.4609, + -0.0016, + 0.0004, + 1.0002, + 0.0004, + 8.064, + 0.0914 + ], + [ + 11.2674, + -0.002, + 0.0002, + 1.0002, + 0.0002, + 8.0732, + -0.0142 + ], + [ + 12.0751, + -0.0002, + 0.0009, + 1.0003, + 0.0009, + 8.0718, + 0.2182 + ], + [ + 12.8778, + 0.0009, + 0.0008, + 1.0005, + 0.0008, + 8.0936, + 0.1081 + ], + [ + 13.6829, + 0.0038, + 0.0015, + 1.0007, + 0.0015, + 8.1044, + 0.1823 + ], + [ + 14.4876, + 0.0044, + 0.0018, + 1.0007, + 0.0018, + 8.1226, + 0.3151 + ], + [ + 15.2961, + 0.0073, + 0.0023, + 1.0007, + 0.0023, + 8.1541, + 0.2434 + ], + [ + 16.1079, + 0.0092, + 0.0022, + 1.0006, + 0.0022, + 8.1785, + 0.2183 + ], + [ + 16.9194, + 0.0097, + 0.0021, + 1.0007, + 0.0021, + 8.2003, + 0.2977 + ], + [ + 17.7247, + 0.012, + 0.0019, + 1.0003, + 0.0019, + 8.2301, + 0.5019 + ], + [ + 18.5499, + 0.0152, + 0.0019, + 1.0005, + 0.0019, + 8.2803, + 0.294 + ], + [ + 19.3613, + 0.018, + 0.0018, + 1.0002, + 0.0018, + 8.3097, + 0.5381 + ], + [ + 20.181, + 0.0202, + 0.0019, + 1.0002, + 0.0019, + 8.3635, + 0.4827 + ], + [ + 21.0109, + 0.0217, + 0.0018, + 1.0002, + 0.0018, + 8.4117, + 0.4352 + ], + [ + 21.8389, + 0.0219, + 0.0013, + 1.0003, + 0.0013, + 8.4553, + 0.5616 + ], + [ + 22.6681, + 0.0248, + 0.0012, + 1.0002, + 0.0012, + 8.5114, + 0.5495 + ], + [ + 23.5035, + 0.0257, + 0.0008, + 1.0001, + 0.0008, + 8.5664, + 0.6455 + ], + [ + 24.3489, + 0.0303, + 0.0013, + 1.0002, + 0.0013, + 8.6309, + 0.562 + ], + [ + 25.1976, + 0.0307, + 0.0015, + 1.0003, + 0.0015, + 8.6871, + 0.5735 + ], + [ + 26.052, + 0.0323, + 0.0015, + 1.0002, + 0.0015, + 8.7445, + 0.6293 + ], + [ + 26.9104, + 0.0347, + 0.0017, + 1.0005, + 0.0017, + 8.8074, + 0.6244 + ], + [ + 27.7751, + 0.037, + 0.0021, + 1.0007, + 0.0021, + 8.8698, + 0.6727 + ], + [ + 28.648, + 0.0421, + 0.0022, + 1.0006, + 0.0022, + 8.9371, + 0.5461 + ], + [ + 29.5212, + 0.0447, + 0.0023, + 1.0008, + 0.0023, + 8.9917, + 0.8209 + ], + [ + 30.4082, + 0.0479, + 0.003, + 1.001, + 0.003, + 9.0738, + 0.585 + ], + [ + 31.2986, + 0.047, + 0.0028, + 1.001, + 0.0028, + 9.1323, + 0.6027 + ], + [ + 32.1931, + 0.0503, + 0.0031, + 1.0009, + 0.0031, + 9.1926, + 0.8186 + ], + [ + 33.0979, + 0.0518, + 0.0032, + 1.0013, + 0.0032, + 9.2744, + 0.7072 + ], + [ + 34.0062, + 0.0534, + 0.0028, + 1.0012, + 0.0028, + 9.3451, + 0.7215 + ], + [ + 34.9247, + 0.0575, + 0.0033, + 1.0012, + 0.0033, + 9.4173, + 0.7555 + ], + [ + 35.8413, + 0.0606, + 0.0032, + 1.0011, + 0.0032, + 9.4928, + 0.7423 + ], + [ + 36.7802, + 0.061, + 0.0034, + 1.0015, + 0.0034, + 9.5671, + 0.4912 + ], + [ + 37.714, + 0.0635, + 0.0032, + 1.0015, + 0.0032, + 9.6162, + 0.774 + ], + [ + 38.6526, + 0.0643, + 0.0027, + 1.0016, + 0.0027, + 9.6936, + 0.8083 + ], + [ + 39.6127, + 0.0679, + 0.003, + 1.002, + 0.003, + 9.7744, + 0.5879 + ], + [ + 40.574, + 0.069, + 0.0024, + 1.002, + 0.0024, + 9.8332, + 0.6322 + ], + [ + 41.54, + 0.0739, + 0.0025, + 1.0023, + 0.0025, + 9.8964, + 0.7491 + ], + [ + 42.5189, + 0.0753, + 0.0022, + 1.0026, + 0.0022, + 9.9714, + 0.4674 + ], + [ + 43.4949, + 0.0774, + 0.0017, + 1.0025, + 0.0017, + 10.0181, + 0.7418 + ], + [ + 44.4731, + 0.0796, + 0.002, + 1.0028, + 0.002, + 10.0923, + 0.7677 + ], + [ + 45.4688, + 0.0837, + 0.0017, + 1.0026, + 0.0017, + 10.169, + 0.583 + ], + [ + 46.4721, + 0.0856, + 0.0019, + 1.0025, + 0.0019, + 10.2274, + 0.5995 + ], + [ + 47.4792, + 0.0883, + 0.0013, + 1.0028, + 0.0013, + 10.2873, + 0.56 + ], + [ + 48.4911, + 0.0938, + 0.0013, + 1.0026, + 0.0013, + 10.3433, + 0.5217 + ], + [ + 49.517, + 0.0969, + 0.0013, + 1.0025, + 0.0013, + 10.3955, + 0.5525 + ], + [ + 50.5334, + 0.0991, + 0.0008, + 1.0021, + 0.0008, + 10.4507, + 0.5985 + ], + [ + 51.5687, + 0.1035, + 0.0005, + 1.002, + 0.0005, + 10.5106, + 0.424 + ], + [ + 52.6083, + 0.1078, + 0.0006, + 1.002, + 0.0006, + 10.553, + 0.5447 + ], + [ + 53.6507, + 0.109, + -0.0002, + 1.0017, + -0.0002, + 10.6074, + 0.3741 + ], + [ + 54.7019, + 0.113, + -0.0004, + 1.0014, + -0.0004, + 10.6448, + 0.3706 + ], + [ + 55.7538, + 0.1194, + -0.0005, + 1.0012, + -0.0005, + 10.6819, + 0.4656 + ], + [ + 56.8074, + 0.1212, + -0.0005, + 1.0012, + -0.0005, + 10.7285, + 0.3772 + ], + [ + 57.8776, + 0.1224, + -0.0012, + 1.0006, + -0.0012, + 10.7662, + 0.2135 + ], + [ + 58.9418, + 0.126, + -0.0008, + 1.0007, + -0.0008, + 10.7875, + 0.4392 + ], + [ + 60.011, + 0.1313, + -0.0007, + 1.0008, + -0.0007, + 10.8314, + 0.2563 + ], + [ + 61.0942, + 0.1372, + -0.0007, + 1.0008, + -0.0007, + 10.8571, + 0.1101 + ], + [ + 62.1665, + 0.14, + -0.0007, + 1.0008, + -0.0007, + 10.8681, + 0.4464 + ], + [ + 63.2474, + 0.1409, + -0.0002, + 1.0007, + -0.0002, + 10.9127, + 0.3852 + ], + [ + 64.3365, + 0.144, + -0.0007, + 1.0007, + -0.0007, + 10.9513, + 0.1574 + ], + [ + 65.4203, + 0.1503, + -0.0001, + 1.0009, + -0.0001, + 10.967, + 0.2386 + ], + [ + 66.5075, + 0.149, + -0.0005, + 1.0001, + -0.0005, + 10.9909, + 0.0 + ], + [ + 67.6069, + 0.154, + -0.0003, + 1.0002, + -0.0003, + 10.9909, + 0.0 + ], + [ + 68.6966, + 0.1604, + 0.0001, + 0.9999, + 0.0001, + 10.9909, + 0.0 + ], + [ + 69.7886, + 0.1624, + 0.0008, + 0.9997, + 0.0008, + 10.9909, + 0.0 + ], + [ + 70.8966, + 0.1674, + 0.001, + 0.9997, + 0.001, + 10.9909, + 0.0 + ], + [ + 72.0083, + 0.1712, + 0.0015, + 0.9997, + 0.0015, + 10.9909, + 0.0 + ], + [ + 73.11, + 0.1757, + 0.0015, + 0.999, + 0.0015, + 10.9909, + 0.0 + ], + [ + 74.2129, + 0.179, + 0.0022, + 0.9989, + 0.0022, + 10.9909, + 0.0 + ] + ], + "turn_indicator": { + "command": 1, + "command_name": "DISABLE", + "keep_selected": true, + "held": false, + "logits": [ + -16.221363067626953, + -6.708434581756592, + -2.558558940887451, + -3.9311578273773193, + 6.994775295257568 + ], + "probabilities": [ + 2.8843449850768366e-10, + 3.903549895767355e-06, + 0.00024758701329119503, + 6.275029591051862e-05, + 0.9995858073234558 + ] + }, + "predicted_agents": { + "format": "npz", + "key": "predicted_agents", + "dtype": "float32", + "shape": [ + 12, + 80, + 5 + ], + "data": "UEsDBC0AAAAAAAAAIQB33n+g//////////8UABQAcHJlZGljdGVkX2FnZW50cy5ucHkBABAAgEsAAAAAAACASwAAAAAAAJNOVU1QWQEAdgB7J2Rlc2NyJzogJzxmNCcsICdmb3J0cmFuX29yZGVyJzogRmFsc2UsICdzaGFwZSc6ICgxMiwgODAsIDUpLCB9ICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAgICAKBGidwERxyUB6Gom8cXt6Pwihh7zc6JfACwzKQC5TszyHTHs/fa+xPMLCkcBCq8pAYj9YPfBefD/Cw1Y9IKGLwBJqy0BTyJQ9bhF9P+H+kz18oYXAGyzMQCB9sD1vwX0/7tKvPagkf8D12MxAJzXCPWYXfj/4oME9YHRzwBJyzUAoQ9I9s0x+P5G/0T1YhmfARBjOQCdp4z02ZH4/9e3iPaTTW8BRqM5Ah5bxPWBffj+IGfE9+P1PwP4Xz0D4m/49FCt+P5AG/j3sokTAt6HPQG4GBz646H0/O6sGPsQOOcDGGtBAdxUQPguEfT8unw8+OMUtwC6B0EBFzBY+MnR9P95RFj7EUiLA9gzRQNxtHT4//Hw/qc8cPmyhF8DdeNFA2MgiPsDXfD9PHyI+kDENwO+10UDkKCQ+2mp8P29cIz7g6QHALknSQA0fJT4JEXw/bjUkPhAz77/4tNJA4e8jPooVfD8VCCM+4FHav1wF00BK3yE+rLh7P8jaID5wAca/klfTQBZJHj47wns/oEkdPrClsb9v2tNALzUaPqSsez/OMRk+uJmcv1co1EBXdRM+Deh7P1+IEj54FYm/6nDUQBHdCz5YCHw/9P4KPiCqaL+MqtRAw+cDPuhqfD/wKQM+8CFBv7Xy1ECOPfc9oex8P9wN9j3whxe/kzLVQEr85j19PH0/hvvlPcAg4L6DYdVAj/PWPZ30fT8yStY94OyPvuKE1UCKzsQ9Nwl+P+8zxD0ATwa+GJ3VQFNPtD14YX4/FtuzPQAQlTx719VAAMGnPVKgfj+9Zac9AIkzPgjj1UA1AJo95N9+Pxi8mT0AxqQ+GA/WQDmziz029H4/AHiLPWBe+D7iLtZAN8N+PQHlfj9kS349wBYgP/k01kB4EGw93M1+PwOUaz2AvUc/SDDWQGq1UT150H4/zERRPaDZcT8+S9ZAzd1APU/Nfj9gc0A9YIOKP5tf1kCuHS49E25+P7CbLT1o7J0/WFvWQAgzFz3EmX4/oM0WPcDNsT9ibdZAcwoFPbArfj+/kwQ9gADEP5081kB3XvA8yiV+PzyE7zw4rtY/TGPWQC5BzzwP830/gW/OPPAg6j+Nc9ZAbC+vPMjgfT9Cd648SHz9P09i1kCbaIY8sv19P1zihTxUDghAT07WQNvkXjzRvn0/dupdPArNEUD4MtZA+XYrPB/jfT88wio8FNUaQEwt1kD/p7s7kZ99PyDJujtMcCRAWP/VQBUYRju7130/ckJFOyZhLUAq99VAqmONuP6VfT8AuYy4HnI3QFK51UBz47i72HV9P8z4t7uGlUBAzKvVQC0HvruPKX0/svm8u9CsSUC3rtVAL+AivHHJfD/b2iG8IJlTQCuA1UBL4GG8bqp8P6NoYLz0VFxAn23VQKMLiLxiL3w/7giHvNL+ZEAYRtVAARyivIn+ez+q2KC8cFduQAwI1UAbt7+89M57P4gnvrw61XdAXOPUQJ9A3rzJk3s/mljcvF5TgEBWxNRAGzD4vN3Xez8YMfa8ZAmFQFtq1ECjNwS9e2h7P/0KA70uk4lAylPUQBhpDr3ni3s/ny8NvTNJjkBl+dNA4awavQa+ez9EaBm9WTGTQLMU1EBhIR69eAt8P6btHL2Uo5dAI+XTQNTEKL07+3s/yXcnvfBmnEAattNApAIzvXuJfD/v0zG959+gQD3a00DiITK9RJZ8Pwr5ML0Bo6VAzb3TQLZ3Qr1SD30/IWNBvVKCqkATedNAckNDvUhffT9PTEK9i7quQDRG00Dx8Ue9ytt9P/slR73pr7NAeDbTQNXbTr2CCn4/bxxOveBGuEBMKdNAynhUvUAxfj/qxFO9Q+W8QB4o00AtzVu90oZ+P7g4W73kX8FAhibTQJsGXb0vzX4/249cvei6xUDEMdNAXGFhvUj2fj/s+mC9cpLKQDxG00BX8GW9LBx/P3aZZb05585AmEXTQNPPar3nNn8/CoRqvQu400C0adNA5LBwvWg0fz/kYnC9Zn/YQOJN00BmhG69il1/P/RJbr1WxNxAykvTQJ12b70kjX8/WFJvve4h4UCzKNNAdEp1vWE7fz8A/3S9WsjlQK1Q00ADOnG9GAl/P4HXcL1EcepAgmLTQFUWb72Wun4/t49uvfR8JkEZ72JAmkcmPWuSfj+01iU98LI6QeTpYUCrWSI9+Lt/P4lJIj3T5E9BXR9hQLHxGT2EO4A/IBoaPToSZUGE82BASiAYPXuWgD8zfhg9CPV5QbrPYEB0QhM95seAP4S5Ez1gRIdBEpRgQN8SCD3LtoA/P3cIPZmXkUHpmGBAcIEBPWKSgD9CzgE9m8SbQXejYED3he08AWOAPxfm7TzGCKZB23FgQP/t2jxcJIA/bhDbPEAwsEFesmBArELKPPQMgD+JT8o8NjK6QeGEX0C6zbQ8ReZ/P4XGtDx7LcRBYrVfQBYdqDww338/0hOoPAwOzkHOnF9An2mGPDsBgD8Ka4Y80MvXQcWwX0DO4XY8lRqAP6P8djxWa+FBN1hfQJwmOTw8WoA/YGg5PEIF60HO315AoeUePN51gD8cLx889Gb0QTPBXkBc77Q7lbSAPxxvtTtQvP1BW3ReQDhQGzszAIE/rOsbO/p9A0Id611AP35cOtM5gT+MjF06HB8IQto9XUCFI3+63WSBP5hDgLrnpwxCVFVcQFYVl7qxmoE/uAeYuk00EUKbeVtAaZm8upW9gT+w4b2616MVQj/aWkBTai67sbGBP9SRL7tbDhpCsOZZQEBfQbu3xYE/AbZCu0RhHkIt81hAVexsu7ixgT/MfW67dqoiQsIDWEB5LX+7nauBP+ZrgLsk+yZCHJJWQEa4k7tJjoE/KZ6Uu4QrK0KedFVALZimu3dpgT9/g6e75EkvQr/+U0C+cMu70UiBPzh2zLsCbDNCW3dSQBXSALw0MIE/VGsBvLd1N0JXzVBAg/MhvKoagT+spiK8LHM7QoQxT0ApwkG84QaBP7WJQrwIZz9CYm5NQOQ0dbx62YA/YAZ2vE9TQ0I6uUtADoCRvBDCgD9W75G81ChHQt3RSUA6cKa8AtSAP4n7przsDUtC8CVIQKcWxbwxu4A/NanFvDDRTkLHIUZAUD3UvOfDgD/H4tS8Zo9SQi80REBO8Pe8GNSAP5XC+LzsOFZCHBJCQIhgBb2ozYA/tc4FvdfPWUJQA0BA8KUVvTDzgD9gOBa9u3RdQjb5PUCcniK9ROuAP4w5I73GBWFCv+g7QGSULb1ZBoE/9EwuvbOAZEJOhzlAkT0+vZQhgT+OHT+9gv5nQhRFN0ALoU29XjyBP0aqTr3nYmtCa2I1QC8pVr3oWoE/9FdXvfzQbkJrgTNAoplcvZlpgT/13l29YCVyQhhOMUBodma95HuBPw/cZ73AinVC5qgvQGyXbr0ddoE/eAVwvZO/eEIzJy5AYE50vZxRgT81o3W9RvF7QvCNLEB/UXO9CFyBP7WudL06NH9CppsqQMlYd72XN4E/L5l4vd8tgULjXyhA9Vd4vWkhgT9IhHm9Mc+CQjRUJ0Dy1nO9fCeBP+sCdb3xWIRCIrklQInac730/4A/3OB0vYjohUKAxyNA1KtvvYr1gD9Jo3C9THKHQpExIkCptmq9K86APy+Ea70O/4hC5EkhQPx7ab3q6YA/k2FqvdR4ikKgyh9A+CxnvbXJgD/h8me90gWMQvjhHkCH7mK9mceAP1+uY70yg41ClKsdQKFMXr2R6YA/eCVfvWz+jkLu1RxAQ99avbfogD+cs1u9yHGQQkUtHED/TlS9JveAPy4oVb1j6pFCcvgaQOhFUb3pFYE/zDRSvVxik0KngxpA6ItNvdg2gT+WkE69FNqUQllTGUAmGUi9NVeBP6cvSb0mO5ZCz+cYQHWdQb1KeYE/F8RCvbmvl0Jq+RdAt8Y7vYmhgT90AT29KRuZQstNF0CmPzi9xcOBP86MOb3khZpC6gUWQJq5Nb1qxoE/3AM3vT7om0JniBVAI3Auvbb7gT/l0C+9YFKdQs4cFEDaazC9QhCCP/DeMb3SvJ5C1ccTQO/aKr24IYI/i00svSwUoEKC4xJANw8ovc4agj8Edym9iHChQnLkEUAzeiG9xSmCP+fcIr340aJCvJYRQLUFHr2lMII/12QfvVorpEKKIBBAJe0bvbklgj/WQB29JYqlQjx5D0Cm9Ra99BOCP78zGL0c4KZCeywOQM4FFL1YE4I/MT0VvVMgqELMPA1A+Q0PvWrbgT9jGxC9wXipQiklDEAp4gy9rr6BP4/bDb3SV3NBi6VTwLDtAD1fdXw/GAwAPSgXgUHjl1PAaEABPVQnfT8wiwA9C8yIQSrUU8Bwq/w8pth9P3Sg+zzwiJBBcadTwBT6Az0qsH4/baYDPVQxmEFnklPAGp3/POdvfz93Wv88kMufQT+SU8Dvq+U8ktd/P6ad5Txyb6dBIWZTwEXr1jzMAIA/GO/WPG4Hr0GTP1PAZza9PHbifz+kLb08gLm2QZZAU8A/Wa88zpp/P004rzwcZL5BU/dSwHUTmzyzVn8/XeGaPO4FxkHlRFTAgiCDPFv2fj8x3YI8Zq7NQSr7U8AIxW08Cdh+P6c8bTyQSNVBjfhTwIJrRjzAzX4/cvVFPPvh3EEy/lPAaAc9PN7Nfj/rljw8yXXkQbBFVMCD0ww89jF/PxCbDDw4C+xB6pBUwIpA4zstQH8/oeviO8GG80F0bFTACKSSOy50fz8MfJI7vP36QRbAVMCSv986AvB/P5i43zrhQAFCRglVwM+qCDouE4A/DLUIOjL8BELzlFXA6PBQul4tgD/wFVG6KK4IQtvzVcDmzSe6pz2AP1D2J7qYbQxCE1FWwKGrD7o4YoA/wOIPut8gEEJByFbAPX3juiZrgD943OO6dM0TQrAWV8DDXcC6YJCAP0LKwLrHdBdCP6ZXwCBo37pFlIA/iOnfut4gG0JfCljA6mCZus+wgD/aypm6L8YeQiWXWMB+Dve5e7eAP5C/97lDYyJCju9YwHe0NLmFqYA/ICw1ucgDJkKcg1nAQ4sGOvK0gD9c6gY6bpwpQsrtWcDCTjA63caAP7jXMDr0Oy1CmmlawDwaqDqH1IA/zKWoOi/LMEL21lrAjijwOiTWgD928fA6L1k0QoRqW8BR0+Y6atCAP0CP5zrC8jdCE8tbwIfTAjvqzYA/xTwDO3F7O0J3Q1zABLAjO2XcgD/1PCQ7CAw/Qv+DXMB4Veo6XtiAP4ob6zpljUJCEMNcwEabVTvwyIA//EJWO+YaRkIQVl3AdxQMO02/gD8ofQw7E5BJQtzjXcAA5hs7M7mAP81WHDtbD01CJidewODrBztov4A/g1EIO92MUEIDvl7AmnbGOuKygD9KAcc6ZwZUQsS8XsD9nBY7CKaAP7D+Fjuwe1dC/ylfwI5eYDo1v4A/JAZhOg7qWkJy5F/AmEFDuNXagD+A6EO4IlleQsYoYMBj6TO6UOSAP9iJNLriyGFCOK9gwNc2hrou84A/VraGur4zZUKa8GDA2sicuuYTgT/UcZ26nK1oQiZJYcBovhu73vuAP6hXHLsvAWxCBn5hwON1WLsKC4E/u1dZu3xlb0Ik6GHASfxVuwgqgT909Va7GNFyQixvYsA4t4C7RSOBP7ZJgbv4K3ZCaTFjwJ0Jj7vjJoE/cK6Pu5yWeUJBXWPAvZemu28vgT9KXae76vN8QuJmZMBp1cS71x6BPxyyxbskKIBCUsJkwASCy7vXJIE/+mrMu9TXgUIjw2XA4/r8u3AZgT9UEf67boeDQk8CZsA/+gO89BKBPy6IBLypLoVCXvNmwErVF7yG8IA/OWQYvNrehkKbXWfAD3wnvDnZgD+MCii8voaIQo9QaMBvLTa8I8iAP1e8Nrx/NIpCYQtpwN3JWbxxpoA/SVhavEbci0J0ZmnATmhnvB6NgD/a6Ge86YSNQr1MasDyZIS8kZKAP3yxhLwgMo9CIvtqwMe6j7zqgIA/GgSQvKbikELt+GvAfyCavFtwgD9OZZq8T4OSQiYLbcB3y568fGqAP8sOn7wKLZRC1JRtwJKbpby4YoA/4dylvNXVlUIt2W7AK+ewvJFCgD/uFrG8IXyXQgcDcMCU1rK8fiWAP5fysryyJZlC8gRxwMxdvbzbH4A/hXe9vA7PmkLIXXLA33jGvP8SgD8Wisa8aXOcQmp9c8ApQsO8neB/P444w7yQEp5CftR0wMubzLwpZX8/o2DMvJe4n0LxBHbA14rHvDgdfz/6NMe8tmOhQsj9dsBOIci81NB+P1itx7zWAaNCb4x4wK0OxLw6i34/UoLDvOqopEK+w3nA6uPCvMPyfT9UHsK8qU6mQl45e8AZbri8usZ9Pwajt7y936dCUJB8wIDvrryKLn0/sfqtvGyBqUJJzX3AOUymvGO4fD/4PKW8zygJwq/VSb3z+QO7W397P9nQArvqngfCNDccvanFYDq+VHw/UilfOvELBsIJWue8t2GCO9myfD+KioE7j24Ewi3fg7zd89s7x1t9P5LR2jtmywLC4P6iu1FiQjxdF34/YqlBPKQZAcIa2kc7c8+FPPiufj8geIU8ks/+wWzbSTzD/Lk8hxZ/P/6puTzkTvvBlYygPLws5zxrUn8/SOLmPAzN98GuusE8o9gPPctFfz8aqA89+ET0wX89+zzRCyM9ryl/PxLNIj3Ur/DBygvcPBAZOz1y634/Ubw6PbYT7cGK4AY92TtPPROSfj8Hs049vGzpwYHzMT2iAmY9GK9+P7h6ZT3CsuXBx0FTPZ+ocT21e34/QANxPXzu4cG6lIA9izx+PUzHfj8ftn093ibewcxNlz2wv4E9bhV/P1ePgT02OdrBqpW0PYL6gj1gQH8/59SCPRJG1sEyl9E9nR2GPfJ9fz/PB4Y9DEDSwWnI8D2cUYY9QJZ/PzBChj0IMM7B/DcJPpM1hj1L3X8/xjiGPQDpycEQ7hk+wniDPebJfz9tdoM9/rbFwV0pLD5f+YA9TeN/Pw79gD2oTsHBhcE+Pj4Zej3i738/RCV6PWYBvcHKLlA+wcFtPXYOgD9J4G09yHW4wfTSYT4GHmQ9dTCAP1BYZD2c57PB/nl1Pp2YWj2qPIA/ttlaPXJVr8GEDoQ+mxtOPc17gD9wik49brqqwQL+jT6gNkA9kJOAP3muQD1k86XBud2UPlkOOD3bnIA/Eoc4PZQ7ocEmnaA+oYwtPUS4gD83EC49dk6cwXJ+qT7tbyY9WuWAP+wKJz14UJfBhKixPm7DGz1kBIE/sWYcPQljksFNPLs+hkgQPXQOgT/K5BA9gWmNwdldxz4hkgg9qgaBP4IhCT0cTIjBNYDNPhAc/DwvCYE/VSb9PBJGg8G0N9c+Qjf0PHQRgT/GQPU8agd8wc1l3D6bGuU8wQyBP/YO5jxciXHBcnjlPhJb0DywAIE/3y7RPNr0ZsEqD+4+wg3IPADfgD+Uvsg8BFZcwbah8j48DLs8ctaAPwKruzyKoVHBqrb7PgWvqzwb2IA/kEGsPPCrRsHKXwA/oqeXPNC0gD/cE5g8yLI7wSa7BT/zA5A8wdyAPxeBkDywtTDBB7oFP/bDgzzQ44A/8zmEPMSvJcGcSAc/S1tvPBbMgD80G3A8rlUawcK/BT+e6E48SdaAP4SWTzzAWw/By8YFP3KnQTwC6IA/iFdCPIb/A8HyygQ/Yf0lPLDcgD/WjCY8QPjwwPK+Az8OFwU8vMaAP5B+BTwQ69nAs90DP+LH5juY2IA/YovnOxDkwsC2GgM/B1nCO5jPgD/G9sI7FOCrwHWcAz9FH347t7GAP8PPfjs4UJXAD80DP8/LKDt3nIA/ADMpO4hse8CHEQI/Upn7OveZgD+oMPw6jMRMwDQ6AT9Pswq1yH6APwD4CrVElx3ApQwBPwkKUrrFcIA/kGZSurAK3b+ikP0+8nh3ujNjgD/Y2He6YAt7v/Yz/D4qFFe7KlqAP/dfV7ug+YG+4h76Pn96jLvHW4A/6qyMuwCP/T6K+fk+1qeluyFUgD9e3qW7wEmePyrO8T5OIMy7Z1KAPy5izLvgpv0/GnDwPvxEALzpYYA/NnYAvNQRL0B0k+8+RVwZvOJsgD/JnRm8Ch5fQHjP6z71Rja8ln6AP5KhNrz8yIdAvOPsPv3XXrzqh4A/Lk9fvAtOoEA5s+Q+rfFkvAuwgD8SkGW8ok64QDCP4D4E5ni8FbyAPxyeebwq9NBAPQDcPjZEjrxawYA/lLCOvFjY6EDuvtk+hNGSvH7SgD8/S5O8E8QAQVZS0j4EHZ6888aAPyeZnrzbeQxBy6jSPpUXrrxkuoA/BZiuvMqyGEFiQc0+b5O3vLS2gD9sGLi8QxclQWDSzD6rssS8td+AP/tgxbwnFTFBgvDJPtJyx7zN5YA/YyjIvAQ+PUGcT8g+gbLLvBTpgD+pbsy8OTRJQQaZxT6KOdG8UxSBP0we0rwoc1VBytG/PhdE2rzLLYE/uEjbvAsuYUEygLo+RlvUvHMygT+KXNW8Gm1tQbQduz51LNm84kqBP25I2rw0dnlBFsu4PrxC2Lzeh4E/ApHZvKjSD8K6V2HAeHQxvAUXez9AwS+8swoOwqzCYMBPSS68tAN8P2ruLLz0PgzCHVZgwC2EGbx3gHw/8ncYvB5pCsKZ0V/A/c77u881fT8OcPq7yY8IwhZwX8C02FG7y/V9P7gCUbsUowbC4U1fwELBPrlKlX4/IDo+uV+4BMKlBl/Al/qpO7cNfz9Aqqk7qb8CwhLjXsBJeR88p1p/PxtGHzzhxgDCcOtewHh9gjzOgX8/A16CPLKO/cGQtF7AZBqlPPeWfz/z+aQ84I/5wefjX8Aw1dQ83Jt/P5+u1DweevXBfrFfwN/K/Tyxdn8/AIz9PG5b8cEDJ1/AAhEZPZS6fz/QABk96iXtwRrXXsBa+yc95Zl/P+DfJz2o6+jBYDRewEnbOD0J8H8/jt04PXai5MEOjV3AG3pCPb8egD/SmkI9jjPgwbuzXMDMbEs9NyuAP9qZSz1QxtvBT9RbwHymVT1CP4A/sudVPbQ/18EDvlrAlA1cPcNIgD+xWVw9tLLSwVyYWcAcrWI9HlmAP9oKYz0w983BBVpYwETsYz1ZQ4A/TjdkPR5IycEgF1fAZJBhPaw8gD921GE99G3EwS7BVcCsh149uTKAP8rBXj2epr/BG4xUwAM0Vz0hP4A/xXVXPeCyusF+UVPA2iBPPfRFgD/EZE89Fqm1wdzfUcCWrEk9gkWAP8rtST3MrrDBxJRQwBvkPT19dYA/+kM+PSulq8EcPE/A/mgwPX9+gD8oxzA9OXSmwUhBTsCTqyk9FIWAPwAKKj1cUaHBUsRMwC5VHT0XmYA/O7gdPWAAnMFReUvAFOMVPXW3gD/JUhY94pqWwSmJSsASLAk9BM2APzedCT29S5HBBjZJwCy1+TyY34A/PZT6PEvni8FK1kfAtGTlPHrZgD9tK+Y8EHaGweISR8AsP8w8xuCAPzn1zDw2B4HBOwJGwGXjvjy19YA/1Jy/PBbXdsFSZkXAzhitPHfygD9ovq08iJlrwcx8RMCJ7pI8hPaAPwl9kzygVWDBmqFDwC2HhDxb4oA/Gv2EPIDwVMEAD0PAs0dwPJvpgD8RJHE8RppJwRIlQsB+1008xfmAPwWhTjyM9z3BMudBwFyXIzzy3oA/LSYkPJJCMsEWpkDALzUiPKcWgT8W5iI86KQmwTb+QMBFjAU8Bh2BPyYhBjyo9BrBs/lAwPcb6Dv9CoE/SQ7pOwD/DsETy0HA/8y9O9kdgT8Qob47llIDwUr5QcDJf6s7pyuBP6FIrDu8ve7AvIlCwM2qkDu8G4E/NEuROyB61sAcA0PAYqA5OwQCgT+BWzo7wFK+wOs8Q8CU3is7bheBPzSaLDuQFKbASoNDwJp6KDubB4E/HCgpO6TgjcDbgkPAMOvEOqXogD8mnsU6ELRswMBtQ8Dt7Ew6ib+AP0CGTTqQKDvAnxpEwG8bYDp8s4A/kLhgOuB1CsCqIETAOAjatyKfgD/Aj9q3wKCyv9NeRMBfcik5gIKAP8DIKTnQeCK/z/JEwD3DPzpbbYA/KBVAOoCXGD4g4kTAhiUMuqBkgD+cXAy6YBJpP0H9RMALNCm6q2eAP5B4KbrAetc/SvxEwCRRZbrfZ4A/MK5lutxKHUADxEXAI9GwukZlgD8YF7G6UOtOQISgRcACsA67zHiAP1rzDrtlb4BAdpRFwJOlU7tCioA/7RdUu2I8mUDh2UXAYOCTu0eRgD9cNJS79GKyQCNdRcAGE9q7lKOAP5Oe2ruqmstA5xZGwJzryLuY2oA/VZfJuxKt5ED5PkbAoarhu47ogD/bd+K7ITb+QEZRRsASDQi8wvOAP9GOCLxHawtBmh9GwJ+MBLzCDIE/9hcFvEdBGEE/lkbAgX0RvBMJgT9mFBK8am8kQbTeRcBKqCe88waBP+BUKLzE9jBB2F1GwNTsNrx0CIE/SKo3vGzDPUH3u0XAUpFDvCgvgT+CeUS86DtKQbqLRcBlXUC8XSqBPyk+QbwytFZBvStFwFkSPryWKoE/lvA+vBclY0FqxkTAiDE6vI5KgT93Iju81MJvQarCRMAy50G8LFaBP/PqQrzt7HtBYJtEwFPiLbxmT4E/j8YuvA8/hEEfyEPAIIsuvBZigT/4fC+8t3KKQURQQ8DinyW8OI2BPzyhJryC3RjCD6/EQFD0zLpoA38/No/MuobrFMKvs8RAXydxO1MZgD9MP3E7MNoQwv+yxEAbpww88JCAP/b2DDx0uwzCCtjEQOoPTzzSzYA/GbdPPCqUCMLaD8VAUXKHPObPgD8b4Yc8umUEwjgpxUCEqZU8NKqAPxYOljy3JgDCLWjFQH5krTzfgIA/cL2tPNTY98FRwMVAhzC5PDpXgD+mcbk8PEbvwb/dxUAqlcI8ZC6AP8W6wjwKwObBAkfGQMJExzwyNYA/r3DHPMQx3sEmP8ZAq+LHPPMUgD+R9cc8vqPVwfmYxkAtUsk8bvd/P2ZRyTxMD83Bcq7GQJXrvDzi5X8/F+S8PMCKxMFk/sZA4mK7PDzHfz8zULs8pgm8wZQNx0BdU6Y8K9l/PzZIpjyEn7PBMwnHQHR2mTxqtn8/jGGZPA0bq8FWNMdAv7WIPD7Bfz/NpYg8EriiweU4x0AlyHU8ogmAP5PSdTwXTprBSEPHQCpgZjzwGoA/YXlmPPjekcEcQcdARjZTPIg0gD9dYlM8koKJwU4Jx0AL4jg8NEiAP7EWOTwfN4HByAfHQJo7Nzx9aIA/4YY3PNjSccFhFMdAZPQiPExvgD+VOyM8KlZhwZbQxkDtHfk7z4uAP0im+TuA2lDBqdrGQN1BzzuckIA/HrfPO2JfQMGEr8ZA86+2O2qGgD/+D7c7kPEvwZiDxkB3K5478W6APxdwnjuEuB/BsWXGQIA2hzujYIA/l2mHO0hbD8FSHsZAmUQ7OxlNgD8HfTs7oEH+wL/zxUC1O8Y6hhSAP5xLxjqY6d3ALbPFQHCmk7e2AoA/AKiTtyRSvcDdNMVA2P2vuozOfz/a7K+6KKidwAwRxUB5ih+7KYF/P/hiH7swb3vASdjEQBObnrupOn8/Bl6eu7RfOsA6bMRAB+C9u+kqfz8kkb27cNb1v35CxEATgvC7Pgl/P3AO8LsA12y/M6rDQNIQHLwxB38/SMUbvICCgT1cbMNA4fpFvB4Tfz/mn0W8cPuGP0z0wkBsRlq8XRp/P1nlWbysvgJAVXXCQJiZfLwKan8/41B8vOo3QkDo6MFARWCQvC10fz/MOZC8UcqAQEmRwUBOVJu8XeF/PzRMm7wGS6BAVDTBQB+KqbzwO4A/XrOpvDMGwEAVssBAdPu4vNBegD/7Qbm8yHffQFFIwEALVMG8mWuAP5qnwbyYBP9A662/QLKj07ygmoA/jCbUvC0VD0EKKr9A05jSvGTFgD8vPtO8WfUeQXOvvkCMsdq8OrCAP21L27z8oi5BhB++QOse5rzaooA/MbXmvF5kPkEk3L1AuyLYvMKugD9+udi8mANOQQtovUBgM928F3uAPy6h3bwgrF1BDfq8QAi54bxhWIA/ngrivOZEbUG4j7xAMrnWvFozgD9r59a8Xyh9QQg6vECaP9G8QzeAP7Bv0bwfeIZBcIW7QASGzbwlLYA/Ba3NvNZPjkFLe7tA92jIvIoEgD8Ub8i89gyWQTfwukALcLq8F+h/P2VpurxY4J1Bnau6QE7NubyT6X8/Nce5vHSspUEkXbpAeYq0vD8DgD+ijrS8mnKtQUYUukD8P6y8SQKAPyZDrLx+OrVBNI25QE/xqrzEIYA/cgmrvN7ovEHeh7lAA3ygvJRFgD/yqKC8LKbEQY9XuUCcBp+85nSAP4JQn7zGdMxBWxS5QDhSn7yYnoA/NbafvGwv1EGXCrlA9C6fvEi2gD+VoZ+80MzbQah+uEDIW5W8y+WAP+rilbx9cONB2GG4QLjjkrzY94A/8XKTvOpJ60GA+LdAEaKWvPI0gT/yWJe8osbyQVyst0C7loW8NTCBPz02hrxyPfpBfia3QADlhrwbE4E/v3aHvOHuAEK9DrdAW96HvBkUgT+xcYi8nLYEQta2tkD+kIa8BwiBP44ch7yMhAhCXPq2QNoteLzVFYE/ajx5vDREDEJftrZAT1JtvEHwgD8ZMm68Ze0PQnDmtkADME688K+AP2q+TryvrBNCeK+2QJAlSbyKjIA/o5RJvBJ0F0JsdbZAts86vJ9vgD+vITu8fhobQlUntkBnnTW8+E6AP+fVNbwWvh5C8wu2QJmtLbxpFIA/3LstvDSOIkKM2rVANQEUvEcNgD8kCRS8RUgzQhJB2EAZ1Ri7gQ99P3/0F7v0cDZCpg7XQLaLlbthGn4/8P2Uuye2OUKb2dVAZ8XXu0nKfj8UQ9e7+vY8QmHU1EBUDe27pCx/P7yr7LuOOUBCfNbTQESX/7tqRn8/9Dr/uzdyQ0KEztJAGC0/vD7Vfj8Yvj68RbhGQqvh0UBZgGm8kJR+P5zbaLxl8klCkOTQQG9ZkbyOTn4/XN+QvBg8TUK51c9Ab1WpvCYofj/ruqi8TXxQQpv6zkBSS828bVR+P5+izLwUuVNCNvDNQAAQ4rwZd34/L2bhvGjyVkIiBc1AipX0vP3jfj+CEvS8ZS1aQsobzEDSvAK9thJ/PxGDAr3CXl1CUiDLQPdGA73ELH8/rRMDvaKSYEK8JspAYxoLva9wfz/e9gq9Vb1jQgoKyUBWtRC9DHF/P8iQEL2z3GZCyhzIQG8AG70db38/TdkavXYFakIU8MZAbcUmvSDAfz+Etia9iCFtQqrexUBhey+9Rs9/P4xxL70ZTHBCsrPEQL1JNb1n/H8/C1A1vbFec0JMgMNAxjo9vQACgD/fRD29hHZ2QpQ5wkBN7US9OyqAP4EXRb1kiXlC2BXBQHJvTb2GMoA/BqNNvbWZfELfxL9AepVUveh0gD/MAlW9dqJ/QrydvkC5AWC9XoaAP6CFYL0EV4FCKEC9QFGdZL1Wt4A/SFBlvS7agkIj37tAnyhnvWLEgD+16We9HVmEQseuukBcMmm9LtKAPwACar3p3oVCuFe5QNX1bL1b4YA/addtvelfh0JmBrhAkxJxvfDmgD/u/XG9PNyIQk5/tkB05nK9weWAP77Sc71+WopCdiW1QPPPeL2O6oA/lcd5veTXi0Kbu7NAuah6vW/bgD+uk3u9+FqNQjZbskDaYHy9HN2AP1pPfb2z045CT/CwQFzIfL162IA/u7J9vdhWkEIYoa9AxeOAvfDSgD/nWIG9QtSRQggXrkCTToC9xdCAP/7BgL0qUZNCQKSsQP8+hL1Kw4A/r6+EvSjGlEI1UatA04aEvUXBgD/B9oS9/UaWQhXiqUC9u4e9a9uAP9g8iL0qwZdCpD6oQATKiL330oA/z0eJvXg/mUJSIKdAKEWJvUTXgD/KxYm97bmaQsihpUCOM4y9MtaAP+62jL3MMZxC3wmkQGS1kb0q5YA/n0eSvcWvnUJbrKJArQCVvZvcgD/4kZW9xSufQrkToUCn1Za9i+uAP/Zxl73FoaBCp3afQE4CnL0Q9YA/C6ucvT0jokKq3p1AG/GevdXjgD8Rk5+995ajQlxlnED4VqS9L72APxLnpL1VEaVCmdaaQNKFpb1dyIA/hR6mvf6JpkKmdplAv0qqvfWqgD+r1aq95wCoQsrul0ARLa69WqKAP322rr2jgqlCMHuWQGWSr70CkYA/dBGwvbD8qkI9G5VAj4KwvXKIgD+w/LC92XmsQsKIk0DQi7K9hXyAP6//sr387a1CnSaSQFs8s738XYA/gpuzvUBpr0K9mJBAp8G0vc1mgD9ZKLW9/dqwQkknj0Dxfbe9g0+AP2vWt71YVrJCLqyNQLLGt72eZIA/ki64vafHs0I+LoxA3zC3vWFIgD/+g7e9aES1QtHZikDs2rm9ZGaAP/1Fur0ttrZCNmGJQM0Jur0SXoA//W66vSguuEKu24dAjp66vVJ4gD9pF7u9k625QrGPhkD0G7u9w6CAP+Kyu72+JrtCJhqFQNivu70iwIA/dl68vVSSvEJMsINA7d63vUXbgD8qnLi9KAy+Qn2HgkCalLa9Tt2AP4ZRt70kh79CATOBQE/ltb3i/YA/dLi2vd70wEIbO39A3kO0vR//gD9yFbW9YmTCQpd/fED2G7K9TyyBP8oJs71o3cNCqMp5QNrlsL0NGIE/qsOxvUBQxUIcbndAIByuvRYxgT+cBq+92sjGQtSXdEDSL6y9tkKBP/0irb3DM8hCpeFxQIvEqL3qOoE/vKypvQKtyUJY4G9ATEGjvS5DgT+mJaS9OhnLQoL3bEDt0qK94T6BP9Kzo722ksxCAMxqQHAFn72EGIE/PsifvbcAzkKXimhAQJ+bvQMsgT/laJy95mvPQttuZkBH9Ze9KwKBP3OgmL0B4tBCsvhjQDMak72P24A/naiTvSSnXEJReXQ9VoelPCzSez/kLqQ8xX5eQol0jD36row8IGR9P1j4izxKXGBCi5SaPeg3YTx0R34/BndgPFg6YkKoLrM9A5dcPF7tfj+KIVw8PSRkQqtHyz0ammo8mFR/P5ZMajx/DWZCmPzhPeVqeDyiYH8/yh54PFEDaEKd8Pk9U8OPPClEfz+Gj488KPlpQtUmBT5Iy6Q8ORl/P2yCpDxo/mtC0+oLPu3vsjwE134/84myPFEFbkK7dBY+Jtu6POyQfj9AV7o8cBJwQoL2DT5Dx8A8YVx+P4krwDytJHJCbjkVPmlUyDwbY34/aLXHPPg9dEKz8hs+WS3JPBhJfj94g8g8Glt2QprSGz7IRsk8jVZ+Px2iyDxAiHhCv9odPiUoxjyXu34/D63FPDqyekKmBR0+5Vy1POwUfz+EC7U84d98Qm5jIT5O6aE8GWt/P5G7oTxyG39CcWkcPonqizzu5X8/SOSLPG6sgELRbhk+UQ9pPE0igD+NL2k8eM6BQnQhEz7XZ0c8pV6APzGyRzxw8YJCblgQPlqKHTxPfYA/x9cdPKUWhEJomAg+kn72O+ObgD/4FPc7HUWFQlc5AD5Ddq073MKAP2b6rTvscYZCOpbyPVrDXDtP5YA/J4ldO0Glh0IJIuM9ZWmbOr7qgD/o95s6UNqIQg2wyT2qcCe5y/GAP9AOKLkIEopChqi0PXQP1rq28YA/lNnWusZJi0JCPqE9iFkhu6vcgD+i5CG76Y2MQg5Wjj0vjUe7cNWAP5kzSLsvzo1CReGIPeFpULtA14A/KhlRu8EVj0Jeplc96AuCu0vFgD8scIK7K12QQruAST3BTYa7Zr2APyqxhrsLqZFCJyEkPeqhU7sCrYA//DBUuxT6kkLtchk9b/L5uiKxgD9mn/q6e0+UQvUPAT0wRPy6SZuAPzrd/LqcnpVCMpLxPKD+hbrGi4A/ykeGulf4lkK8xsU8XHBbOimHgD845Fs6u1OYQubwzzzbPfs6lmmAP4Cl+zrAq5lC9WjAPI4PZzvIaoA//29nOz0Om0IDItA8ZPCpO5lugD/oOao76G+cQlovxTxtsd07L3qAP3Qb3jv5051Cv+AAPY6iETyJcoA/9eMRPLs3n0Jjvh09KPk7PJV5gD/2Ujw8yJqgQsb0CD1TtUo8k4uAP4IkSzxKB6JCRy0kPUIhYTyJl4A/bqdhPD94o0JqeTQ95faAPMW5gD8qVYE8eOKkQqmiPT09sI08wL+AP0YbjjwgWaZC9HtUPXyLlzyl14A/QAyYPG7Hp0Ij83c9ewylPL7cgD87nKU8oT6pQrZUez3PJbM87/2AP1jZszxJt6pCsoeMPUEuujwcJYE/fAW7PPUqrEKgpYw9ACfBPIgzgT9WEcI8AKatQucxqT16KMo8YkCBPx4oyzxWH69CUeOxPVs82jyMQYE/ylHbPBycsEJZS8894hfhPJlPgT+cQuI8GxiyQmJw3T3Pe+g8pT6BPzKh6Tx2mLNCe3b/PS4q8jxNVoE/hXLzPL8StUKKvQM+4zHvPDJFgT8bZvA8UJG2Qn8wET5iq+88qF2BPyL38DwlErhCRfcbPtF+/DzbVIE/JtT9PCKduUJy1SU+vl/1PDdegT8ktPY8ox+7QvIkMT63rPM86EyBPzfu9DzZqbxCLNw8PjbC8Tw+VoE/7gnzPDQzvkJuqks+0FPvPCdegT+Ln/A8OL+/QhA5WT6Hz+s8/FaBP6cP7TzwRMFChJVlPq4r8jxMUYE/S2/zPA7VwkK4FH4+kN7uPFRYgT81JPA8r2PEQgdwhD4QTeo8AjWBP/1r6zwi8sVCk2+KPpk65TzUGYE/zTrmPIiDx0JxbpQ+gtrgPEj6gD/3ueE8YhTJQh3xmz7mE9c8O9mAP5TN1zzOpMpCFV2mPsLJ2jwjvoA/mG/bPPI3zEJ+96w+9yLRPLySgD/BndE8RM3NQgOjtz5wW848Hm6APwC3zjyKb89CKeTAPvgF0DybTYA/5UfQPEz90EKYMs8+j23KPC4qgD+Nkco8pJrSQtDU1z7knsU8ksl/P1aMxTyjMdRC/mPjPn5uyTw3rH8/IFDJPNbL1ULfX/E+7ibCPNRafz+f6sE8KmbXQpIL/z7pIMU8jux+P0q5xDxgel1C6FFpQDGCyTyME34/+MLIPOrFYELcZGhA15GyPMEffz9wRbI8jDNkQuOTZ0AUGpg8saJ/P3r/lzyMmWdCElBnQHvfljxh638/f9qWPJb9akJlDmdA3XibPIX/fz/peZs8sFVuQonAZkCke408/41/PwldjTwZt3FCyYFmQNMwgTxmKX8/XPuAPDQJdUIcP2ZAD0ZWPNatfj9RuVU8UmR4QtLRZUAptSg8TTN+P70dKDxWuntCLM5lQEGZ6jtkAH4/F6/pO8v+fkLKJWRAvDeKO8/YfT/9ook7wiSBQsLxY0BBGjM7mwp+P+RqMjv4w4JC4qFjQEbckzguLn4/wFWTOANghEKQHWNASZRAuqxnfj+0+j+66viFQnmGYkCy7y27Yc1+P46HLbsNjYdCQMdhQCEomLvSGH8/f+OXu38XiUKJT2FA4tz/uzmAfz9cnf+7m6SKQq5PYEDYXTi8zv5/P+ldOLyYKoxCz5VfQNMEYLwaK4A/bytgvASyjUINb15AbEJ/vOtkgD9fqH+8Yi6PQnxkXUD3Co689ICAP25TjrysqZBCYRhcQLwnnLwdpYA/qo2cvPsfkkKx9FpAzySrvIWkgD9klKu81pOTQjWCWUBFrLG8fM+APw8+srxn/5RCDktYQBSBxLyOwYA/EhjFvO9plkJj1FZABsHJvKzVgD8KbMq8DNSXQl8vVUBcU8+8ssKAP+Hzz7ySMplCw+RTQAc00Ly+uYA/+M3QvLmSmkJwTlJAkRPSvKe+gD/5stK8i+ubQnMGUUCu09C8cbyAP01w0bxMP51CjitPQJ0q1rwIuYA/isjWvOWSnkJGvk1AVH/cvPykgD/ZEN281OGfQrH+S0A5HOC8PZeAPzWk4LyJL6FCqMZKQPeL4bxXn4A/AhzivF9zokLTNElA6AbgvBecgD8Wk+C8nLyjQoYFSEAtaem8i5mAPzn56byu+qRC749GQH0R37y/nYA/e57fvD80pkLXTkVACE/ovMuagD9/3+i8MWmnQtzzQ0Czu+K8uoyAPw0847xdnKhCl/pCQNjg5bzHp4A/YHvmvG7OqULnjkFATaHjvNqigD/dNeS8GfmqQrjeQEAPKOK8pZqAP1y04rxaHaxCzu8/QAil37wzl4A/ryzgvCtArUL6sD5A0WPqvCmtgD93Buu8ml2uQqswPkCG9Oa8NqGAP+SJ57yjfK9C0kU9QLxp47yvroA/qQjkvM2NsEKDaTxAs/7jvN+5gD8EqOS8u6WxQvKaO0DCr928C6+AP9BK3rzSsrJCQBM7QOE95bw+lYA/XMflvIO5s0L0lTpAeyHXvMKYgD8Gpde82MW0QsfsOUBoadm8wbaAP+QH2rzmx7VC5rQ5QJ1T37x2r4A/O/DfvE/LtkJ7ajlAnBnVvNOigD87pNW84cW3QiFCOUC6gM68AKeAPz4Kz7z4w7hCLMg4QAuQxbwbr4A/oxnGvJS1uULcqjhA3Qq/vAWegD8Cg7+83aS6QraFOEBTprm838WAP9w3urwukrtCtv83QILZw7wgv4A/H27EvGl+vEKIYzhAGTjDvGTWgD/23cO8e2G9QhgkOEAIE768pd2AP8q5vrxMSb5C8S44QCodxLyWD4E/oO/EvBAiv0IEQzhAmKfCvNsKgT/bdMO8MwPAQggIOEButMG8BC2BP4Sawrx25sBCZkw4QLPhxbxeSIE//+HGvCi+wULX3zdAOabFvIp7gT+5zca8UIfCQj04OEDrCru8upaBPy42vLxuX8NCL7k4QBphuLz4roE/gZm5vMozxELYBDlAgd20vPrPgT8zJ7a84vzEQgBeOEDYm7i8Ht6BP6T2ubyAwsVCaLA4QK0es7wxCII/fYy0vE+PxkKnwjhAniy0vIIIgj/UnLW8nVrHQlI7OUAnaq68rDGCP4bqr7x8G8hCVtQ4QGxMrbxbOYI/hM+uvDPXyELtBDlAVr+gvHJFgj/FLaK8TZ3JQkGQOUBSD5G8hFWCP+BikrxhV8pClaY5QGuokLy0UYI/3/iRvLQXy0JbKjpA8X+AvKYugj8JmYG8+sjLQpWIOkBmUXG8wDmCP5lrc7zgc8xCgbM6QHJiUbx3CoI/gg5TvGkxzUIvIztAPYUxvBDhgT9G0zK8vCeFQv+wcD1z+qk8Tgx6PyQCqDzgLIZC7S2LPQyCjjx66ns/9F+NPDc1h0I3uZo9xdFpPBkjfT8UhGg8Dj2IQtoftT05bXM85gJ+P1F8cjzcS4lCkkvPPWTSiDx4in4/Ym+IPMpZikLbQ+k9pT+UPKikfj8b3JM8Cm+LQl7vAT5rPKs8K5p+P1bGqjz0goxCfbsLPmcjwjy2cn4/FI/BPGqejUIZJBQ+fYHQPFpAfj8Qzs88+7mOQhJgID7PW9c8Ww9+PxKO1jz21o9Cj0sYPoTE2Ty1730/F+fYPKz3kELePyE+aw/fPGAEfj/HNd48IBuSQqnGKD6R59s8PvV9P2kK2zzLPpNCdhspPvqv3DxP/n0/99XbPP1qlELUbCs+nwHXPKdUfj9QUdY8SJSVQrWtKj4CasQ8NJ5+P7DkwzxRvpZCUI4uPiu5rjwl9n4/I2CuPL/vl0Ly1Cg+KKaVPEpsfz8MfJU8lCCZQo4qJT7/gnM8gst/Py1rczzCVZpCL64ePl36UDyPIoA/TRdRPKGIm0K30Rs+OcslPFNHgD/H+SU8/b+cQtXIET5Bmfs7XmmAPyAB/Dve/Z1CpekIPjA+tTtZjoA/FqO1O/A7n0Jy+f89gSx3O+y5gD8X4Hc7In2gQv6H7z104NY6Gr6APwiA1zrvv6FC7h7UPfWdtzkUxIA/mCq4OeYFo0Kpf7w9p/RnumK+gD8ooWi6zkqkQvVKpz1vA8a6S6mAP2SGxrq/maVCKZuUPYUQ2rpmpIA/kpzauh/lpkIXqY494bqtuoingD+ULK66xzaoQmWjXT2eh+K6V5KAPyAJ47pliKlCTn1QPdqn4Lp0hIA/GhzhuoPdqkJXuyw99h6nuXp0gD8Aa6e5PjesQr6ZJj2TesI6JHiAP9rVwjrkkq1CD0INPX1qADtfX4A/V5oAO4HrrkK/Bwg9Zl9jO3RRgD/Np2M7pkmwQvb7+jztcr47fkuAPzqrvjvtqrFC/RUIPfIJ6jvaK4A/SjLqO78Js0J1Mgc9Fh4XPNkqgD+nNxc8fnC0QorVGj3KVDw8li+AP1R4PDzD17VCa4ogPawFXDxeNYA/YzRcPJQ/t0KW3Uk9cm+DPCIygD/oiYM8Zaa4QsS1ej1/qZ48uDiAP+vNnjzTDrpCJ515PTHhqTxoR4A/IxKqPNZ8u0J43JI9z6K6PFpZgD8F5ro8AfG8Qk2QpD1fts48nHuAP/8czzy4Xr5CYLm2Pa/14DwMiYA/vnHhPIfXv0JgdNA9xJzxPE6kgD9VPPI81EnBQrl+6z1Lkv48EK2AP6ZD/zwMwMJC/yX8PSVqCj0dyoA/ztoKPcw8xEK+LQs+uLkPPdH4gD8zSRA9MbPFQsqTEz7UuRM90QSBP3NUFD1IL8dC/QkpPu5dGj0yEIE/wgYbPXiryEJI2DQ+pnwjPdMZgT81NiQ9XyrKQjJNSj6N2Cc9yCmBP9WhKD2kqctCajBaPqBNLT14HoE/NBYuPdspzULtKnE+0GAzPUE6gT9iRDQ9uaXOQiacez7sjTI9HyiBP7xjMz14JtBCHv+IPuunMj36QIE/NY8zPY+n0ULWKpE+Frk4PbY8gT+spTk9ojTTQr5tmj6abDU9iU2BP5tgNj0Lt9RCn9SkPu/WND0kOYE/srs1PT1C1kKafK0+atAzPcpIgT/JvjQ9pMzXQoXTuT6XPTI9JlSBP6gxMz3HWtlCXU3DPjJBLz2dU4E/kjAwPVzg2kJ2Pc0+RRYzPZVPgT9fCDQ9ZHDcQqs73T5o0zA9/luBP9jKMT0o/91CuNrmPqEhLj2UOYE/rP0uPbKN30LwS/A+fIgqPZ0lgT9pUis91RzhQui//T6Yjyc9VAaBP05BKD0Fr+JCF3YEPwjPIj3G6YA/NmkjPXpA5ELenAs/v4kiPbzPgD8eEyM9mNPlQqVEED9aGh09OaaAP06FHT1uaOdC3eMXPxfJGz3fhYA/YB8cPQ4J6ULk6B0/vZkbPQtngD8r3Rs9B5XqQk7yJj92bBc9aj6AP8yVFz0lM+xCa6EsP8FJFT3g538/80YVPcbG7UJJvDM/IisVPdu6fz81GxU9C17vQpJIPD+wYRI9ZGl/P506Ej2s+fBCBK9EP9SWEz3M6X4/t0oTPWA9kkJLgmDA5zV+vFxbez8I6Xu8oMuTQuquYMDTdYi8TXl9P0TKh7yqXZVCAdNgwKTPdrwV+X4/GFJ2vCztlkISdmDAl0wUvCfVfz9wQBS8M3+YQgIbYMDMy0a7ixmAP6zfRrssEZpC/blfwI56eTqRCIA/6IJ5Ok+nm0KoS1/AZmDLOz/Nfz9oTMs7QDedQmrQXsB2Jhw8+F5/P6f1Gzzzy55Co3RewK5ZTDyD8X4/Zu5LPJRfoEI5u13Aq2hxPNymfj8Mx3A8i+2hQpa9XsD0PoQ8snF+P8/YgzztgaNCGQNewJoMmDygqH4/vqeXPFcRpULed13ALr2iPH60fj8tVaI8mp+mQt0wXcCeErw8Yeh+PwSuuzyeMahCEsxcwCpqxDwjPX8/0SHEPGm+qULCZ1zAAFjJPCtnfz9+Hsk8kEarQuDCW8ApLco8o6l/P68NyjzM0axC1plbwFiswzxE8n8/eqnDPBJbrkLyMFvAE9/BPKUJgD+y6ME8AuWvQqIAW8CElso8oxyAP9KvyjwwZbFCtFpawFjm0jwFHYA/PAHTPIbtskIQJFrA3sbWPPgqgD8S7tY8nm+0QrrUWcAyC9w85SeAP+Aw3DxE8rVC0aNZwFaF6DwSP4A/n8LoPDpvt0IiNFnArh7qPFcwgD/4Tuo8me64Qvf8WMCIJ/E8XjOAP2Jc8TzgbbpCtcdYwAWU+jyxIYA/Abr6PH3ku0JQgVjAA8kBPewGgD9NzwE9cV69QrQ0WMBCDgg95QeAP6gVCD2u1L5C95NXwAEMDT33DYA/RBcNPT9LwEIUlFfA5T8QPaoDgD/HRRA9csDBQiRDV8DPQRI9u8Z/P201Ej26MMNCgvFWwIMrFj3hsH8/nBgWPSumxEJlbVbAs8IbPfqrfz/xrRs98xPGQu8lVsA7LB49Ppl/P4QRHj01hsdC0cZVwAAwHj1bqH8/9BkePc/uyELIIVXA9+AkPRiSfz9EwyQ95VnKQo/IVMB/8SM9zmZ/PwrGIz20wstCC4lUwIbdKD0RXX8/560oPU4vzUJps1PAnZMrPRR7fz99bSs9+ZbOQvJ2U8C+xyw9v3J/P6GeLD3r/M9CQJ9SwKPdMT11Vn8/46kxPfBg0UK31lHA4Nw2PQt9fz/ftTY9EMPSQmOmUcBGujU9mpJ/PxKbNT3kJ9RC5P5QwPNeNz2Ps38/aUs3PfCK1UIYjVDAgdI5PaDtfz//0zk9SO7WQrzJT8DQ0z09JguAP8jkPT08UdhC8vxOwIPLQT0OIIA/Cu1BPQ6v2UIjX07AEpM9PUwogD+VuT093gvbQhi2TcDGt0Q9eD2AP7PwRD3IcdxC/0VNwEo4RD0UaoA/OJNEPabQ3UJJTUzAobo/PaV9gD+2IUA9sjDfQq6HS8Bel0A9X4qAP5IIQT3mjeBCFPVKwHauQT2Km4A/Zi1CPZzu4UL3KErARHQ/PcW4gD9mB0A91E7jQsmKScDu3T49DcyAP/B+Pz2yo+RCyNJIwGArPj1f+IA/qOw+PTj95UJCjUjAtIY3PUz5gD9RQTg9CFvnQgKWR8BW6zA9DRGBPxyvMT0ArehCPj9HwI14LT2NIYE/bUMuPZcK6kLdYEbAU2kkPYVMgT+PRCU9RVvrQqd/RcBbvR89RkOBP0iMID2ktuxCdwlFwMCyGD0nZIE/vYsZPZwQ7kJaIUTAeucRPWh9gT/UxBI9PGvvQk3jQ8CntAo9gpqBP32WCz2Os/BCGUZDwFsDDD19oYE/NOsMPaQL8kImMkLAHMIGPZKygT8Aqgc9smLzQtqPQcDaIgQ91KGBP3j9BD2EtPRClmRBwPG0+jzInIE/O078PLwA9kLrfEDASQH1PG+RgT8vhvY8Lln3QjXnP8AjH/E8eoqBPyyX8jxKqfhCrhI/wJYq6zyWe4E/cIvsPPT0+UJapD7AnGDmPAxNgT87kOc8q0X7QvSePcDsAus8/DSBP7gi7DzpmvxCj/o8wI997DzOI4E/WY/tPDnh/UKd3TvAO4LpPELrgD/iXOo8BDb/QmMxO8B26fQ8CJWAP7h89TwWQQBDUak6wHvt+zwKcoA/zGL8PEPgAEP8YjnAwToEPSgzgD8fWAQ9nIgBQ7I8OMBUTw49Vq5/P0s8Dj0BHp9CxihgwCc9YLwNxHo/MPNdvCCwoEJ9X2DAkaV7vOkefT+MPHq80UWiQnKHYMBdX2W8t8l+P1DVZLwV2KNCVzBgwMEYAbwUvH8/zQcBvGJtpUI63l/AzcfxuhwQgD8I1/G6dQKnQsyAX8C2DBI7Jv1/P+oLEjsDnKhCMBxfwDzq9Tu8tX8/3Mb1O0UvqkLdqF7AJwY0PPs/fz8ZwzM81MarQn5YXsD7X2Q8ns9+PyjZYzxUXa1CpqVdwHG8hTxJgH4/+ViFPADurkIomV7ALgeRPPtJfj8TjJA8S4WwQubjXcAqSaU8GHh+PxTMpDzzF7JCY2BdwPyRsjwwgH4/7Q2yPIKos0JYIl3AgSbMPESnfj+/n8s8NT21QkC7XMDNltY8m/d+PyAr1jx3zLZCd1pcwECY2zyeG38/qDnbPPVWuEK8u1vAKQHePLlZfz+JvN08sOS5QhSQW8Dw2dg8aJV/Pwiw2Dz+b7tCwSpbwIDH1zyirn8/ZqjXPP77vEKY8VrAeAXiPA7Jfz/i8OE83n2+QpI8WsDbVOw8Ocl/P8Q/7DztB8BCbgVawDkL8Tzo238/rv7wPFCLwUIqqFnADvf4POzNfz+c4/g8FA/DQrB3WcAurwM9kvl/P26wAz06jcRC1vNYwMI9BT132H8/ejYFPb4MxkKBq1jARxAKPcfYfz8MCQo9Lo3HQmdrWMBjpw89+7F/P0SVDz0VBMlCnBRYwFFvFT04en8/g0wVPW19ykJMqVfAlmscPb51fz82Rhw99vLLQqLsVsARPSM9RH5/PzsZIz1yaM1C2OZWwHh9Jz2sZn8/R1EnPQfdzkJKfFbAyPQpPcYffz+VsCk9CkzQQiUJVsAkBy89bgl/P6i5Lj29v9FCDWJVwHEuNj3hBH8/w9w1PTcr00IgB1XA9rM4PYXqfj/cVzg91ZvUQmyTVMDNKTo9P/x+P4vTOT3gAdZChMlTwOaPQT046n4/FjBBPb5q10JnUlPAEbNBPfTAfj+WQ0E9adHYQvj9UsCrfEY9pbV+P4UGRj3ROtpCyQpSwDViSj2O1H4/XPZJPRCg20Iyt1HAG29LPQ/Dfj/b+0o9jQPdQsHHUMDr7VA9tKl+P89tUD3DY95CDtVPwGumVj3XyX4/8DBWPSbE30Lfg0/ArFhVPf/Xfj+n6VQ92SXhQizETsA6q1c9sf5+P5NLVz2/heJCAkROwA+wWT3hMH8/GWVZPWTn40InVU3A+SBePQpgfz9/6V09AEblQshuTMC0tGE97Yx/P5aQYT1GouZC3sVLwJ9uXD0xnH8/Q1FcPSX750KDAkvA2bFjPS6/fz8HpGM9YV/pQip3SsDZNGI9XBCAPwhSYj3UvOpCRWRJwC8lXD3YIYA/3U9cPZMa7EKpjEjAm1BcPRoygD9WiVw9l3XtQjLUR8Cdllw9I0eAP5ThXD1B1O5CLAVHwJ6dWD0hZYA/IwBZPWgz8EJCSEbArrBXPRaBgD86Klg9WYbxQsOfRcAGS1U9NLOAP7XsVT0F3/JCWV9FwCxjTT3AtoA/1wBOPXk79EI/X0TA0DtFPYjQgD9D5kU9h4v1QgYFRMC+2z89a+qAP3GUQD3s5/ZCPhlDwJr1NT1HHoE/x8g2PeI3+EJNLELAACQwPcUYgT8q7DA9oZL5QqTEQcCfBig9HT6BP3rdKD2s6/pCK9JAwHJaID2fYoE/2D0hPfpF/ELVoEDAzWwXPQyFgT9eVxg9eY79QrwTQMCqbxg9N5aBPxRmGT3j5f5Cqgo/wIdGEj3irIE/mD8TPYweAEO9bj7AIxIPPa6jgT9uABA9FscAQ8RNPsD11AY98qaBP9u2Bz0ybQFD0nE9wIS7Az0Wn4E/CZQEPRIZAkNy2DzAtIMBPQObgT9sVgI9iMECQyAEPMAcWfs8gI+BP27m/DwQaAND+Zo7wJZw9jwRZ4E/B8/3PNAPBENEljrAl9z6PEBSgT8YLfw8zroEQy4JOsBn2/o81EOBP8Ud/DzqXQVDoec4wLZJ9zyRC4E/AFH4PBYJBkPRPTjAcswBPSWzgD8RKgI98q4GQ9K6N8AO1gQ9bI6AP/IiBT0CTgdD0G82wBODCz0hUIA/M7ILPSj3B0PAUzXAoe4VPXvefz8Z6RU9UEsBAi0DLQAAAAAAAAAhAHfef6CASwAAgEsAABQAAAAAAAAAAAAAAIABAAAAAHByZWRpY3RlZF9hZ2VudHMubnB5UEsFBgAAAAABAAEAQgAAAMZLAAAAAA==" + } +} diff --git a/code/tt_diffusion_planner/server/__init__.py b/code/tt_diffusion_planner/server/__init__.py new file mode 100644 index 0000000000000000000000000000000000000000..d70bf51d244f8130752fab56815a9c6cafd5bded --- /dev/null +++ b/code/tt_diffusion_planner/server/__init__.py @@ -0,0 +1,5 @@ +# SPDX-License-Identifier: Apache-2.0 +"""HTTP serving of diffusion-planner-p150, thin bindings of the vendored ``ttaw.server``: ``app.py`` (the ASGI app +tt-model +runs with uvicorn), ``client.py`` (stdlib request builder, runnable as a script), ``smoke_test.py`` (stdlib PASS/FAIL +check of a running server: ETH dispatch and the 12x10 grid asserted, agreement with the stored CPU reference).""" diff --git a/code/tt_diffusion_planner/server/app.py b/code/tt_diffusion_planner/server/app.py new file mode 100644 index 0000000000000000000000000000000000000000..f91d9c6d73d52fc0f2f5a8abcb8b4e09b818a495 --- /dev/null +++ b/code/tt_diffusion_planner/server/app.py @@ -0,0 +1,51 @@ +# SPDX-License-Identifier: Apache-2.0 +"""ASGI serving app of diffusion-planner-p150 on one Tenstorrent Blackhole p150: the HTTP contract of the Autoware +collection (vendored ``ttaw.server.app``, BUNDLE_CONVENTIONS.md section 7) bound to :class:`DiffusionPlanner`. + +Served by tt-model-manager as ``kind: tt-dit-server``:: + + python -m uvicorn --host 0.0.0.0 --port

--lifespan on tt_diffusion_planner.server.app:app + +Routes: ``GET /``, ``/health`` and ``/v1/health`` (always 200: ``ok`` / ``starting`` / ``error``), ``/info``, +``/v1/models`` (stub), ``POST /predict``; errors 400 / 422 / 503 / 500 (SERVING.md section 3). Everything that +touches the device happens in the lifespan: weights -> device (ETH dispatch, 12x10) -> graph -> trace capture of +every warm-up variant, so uvicorn's ``Application startup complete`` (the line ``tt-model serve`` waits for) means +warm; SIGTERM (``tt-model stop``, 120 s) closes the model under the lock. ``/predict`` calls the Python API, so it +returns exactly what ``model(...)`` returns. Importing this module has no side effects beyond importing fastapi and +pydantic (the image's ``verify:`` imports it without a device); the environment is read in the lifespan only. + +Model-specific request fields (e.g. PointPainting ``rois``): subclass ``PredictRequest``, pass ``request_model=`` and +``decode_extra=`` (which adds the decoded field to the call kwargs) to :class:`ServerSpec`, and list the keyword in +``EXTRA_INPUTS`` of the model class. Host tests swap the model before starting the app: +``app.state.ttaw.model_factory = Stub``; ``app.state.ttaw.predict(request)`` is the route handler itself. +""" +from pathlib import Path + +from .. import __version__ +from ..api import DiffusionPlanner +from ..ttaw.server.app import PredictRequest, ServerSpec, create_app, parse_mesh_shape + +__all__ = ["SPEC", "app", "PredictRequest", "parse_mesh_shape"] + +SPEC = ServerSpec( + model_name=DiffusionPlanner.MODEL_NAME, + env_prefix=DiffusionPlanner.ENV_PREFIX, + model_cls=DiffusionPlanner, + task="ego trajectory planning with a diffusion model (DPM-Solver++ 10 steps), neighbour prediction and a " + "turn-indicator command", + default_weights=DiffusionPlanner.DEFAULT_REPO, + owner="changh95", + io="the Autoware planner tensors (ego and neighbour histories, lanes, route, polygons, line strings, goal, ego " + "shape, turn-indicator history) in, an 8 s ego trajectory, predicted paths of the valid neighbours and a " + "turn-indicator command out", + autoware={"package": "autoware_diffusion_planner", + "path": "planning/autoware_diffusion_planner", + "autoware_universe": "9ceaccf026c31ffc5319bc9eeb4bd7bede0af3fd"}, + source={"repo": "https://huggingface.co/changh95/diffusion-planner-p150", "license": "Apache-2.0"}, + calib_dir=Path(__file__).resolve().parents[1] / "calib", + version=__version__, + description="Diffusion Planner v5.0 (Autoware diffusion_planner) on one Tenstorrent Blackhole p150. " + "Not an OpenAI-compatible API.", +) + +app = create_app(SPEC) diff --git a/code/tt_diffusion_planner/server/client.py b/code/tt_diffusion_planner/server/client.py new file mode 100644 index 0000000000000000000000000000000000000000..f6d49d833ebb9f193618a5cbfdd8e42d766881ee --- /dev/null +++ b/code/tt_diffusion_planner/server/client.py @@ -0,0 +1,22 @@ +#!/usr/bin/env python3 +# SPDX-License-Identifier: Apache-2.0 +"""Build (and optionally send) a ``POST /predict`` request for diffusion-planner-p150. Standard library only. + + # the planner tensors (.npz holding the 15 raw tensors of INPUT_SCHEMA), with an optional runtime param + python3 code/tt_diffusion_planner/server/client.py \ + --inputs code/tt_diffusion_planner/samples/kashiwanoha_dense.npz --param stopping_threshold=0.3 --out req.json + curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json + + # send it directly and print the response + python3 code/tt_diffusion_planner/server/client.py \ + --inputs code/tt_diffusion_planner/samples/kashiwanoha_dense.npz --url http://127.0.0.1:20000 + +The implementation is the vendored ``ttaw/server/client.py`` (C08); this file runs it from the model repository +with any Python 3.9+, without numpy and without installing the package. In Python, use +``tt_diffusion_planner.ttaw.server.client`` (``build_request``, ``post``, ``wait_ready``, ...). +""" +import runpy +from pathlib import Path + +if __name__ == "__main__": + runpy.run_path(str(Path(__file__).resolve().parents[1] / "ttaw" / "server" / "client.py"), run_name="__main__") diff --git a/code/tt_diffusion_planner/server/smoke_test.py b/code/tt_diffusion_planner/server/smoke_test.py new file mode 100644 index 0000000000000000000000000000000000000000..49a39a9171c35dcfae803a61571916582fb82f2e --- /dev/null +++ b/code/tt_diffusion_planner/server/smoke_test.py @@ -0,0 +1,162 @@ +#!/usr/bin/env python3 +# SPDX-License-Identifier: Apache-2.0 +"""Smoke test of a running diffusion-planner-p150 server (standard library only; any Python 3.9+, no numpy). + + python3 code/tt_diffusion_planner/server/smoke_test.py --url http://127.0.0.1:20000 --wait 1800 + # a served container package, one serve profile (what code/scripts/container_smoke.sh runs): + python3 code/tt_diffusion_planner/server/smoke_test.py --url http://127.0.0.1:20000 --wait 600 \ + --manifest /diffusion-planner-p150/tt_kernel_manifest.json [--profile ] \ + --out /tmp/diffusion-planner.json + +Checks (all failures are collected; prints ONE line ``PASS ...`` / ``FAIL ...``, exit code 0 / 1): + * ``/health`` reports ``ok`` (waiting up to ``--wait`` seconds), ``/info`` and ``/v1/models`` answer; + * ``/info`` reports ETH dispatch and the 12x10 grid, the p150 target of every published number (PLAN.md 0.3 + item 6, D14): WORKER dispatch (e.g. an ETH open that fell back) or any other grid FAILS; + * with ``--manifest``: ``/info`` runs what the staged package pins for the serve profile + (``DIFFUSION_PLANNER_DISPATCH``, + ``DIFFUSION_PLANNER_NUM_CQS``, ``DIFFUSION_PLANNER_VARIANT``, the weights revision); + * ``POST /predict`` of the shipped sample (the planner tensors as ``inputs``) answers 200 with the documented fields + and passes this model's output gates (:func:`output_gates`: 80 finite trajectory rows of 7 columns, a valid + turn-indicator command, the predicted neighbour paths); + * the served output agrees with the stored CPU-reference output of the sample within :data:`REFERENCE_GATES`: + ``--reference``, else the first of ``..reference.json``, + ``..reference.json``, ``.reference.json`` next to the sample (skipped when none exists); + * a malformed request answers 400 (not 500). + +The checks themselves are the vendored ``ttaw/server/client.py`` and ``ttaw/server/smoke.py`` (stdlib only, loaded +by path); this file holds the model-specific parts: the sample request, the output gates and their thresholds. +""" +from __future__ import annotations + +import argparse +import json +import math +import sys +import time +from pathlib import Path +from typing import Any, Dict, List, Optional + +HERE = Path(__file__).resolve().parent +sys.path.insert(0, str(HERE.parent / "ttaw" / "server")) +import client as ttaw_client # noqa: E402 (stdlib-only modules of the vendored ttaw) +import smoke as ttaw_smoke # noqa: E402 + +MODEL = "diffusion-planner-p150" +ENV_PREFIX = "DIFFUSION_PLANNER" +EXPECT_DISPATCH, EXPECT_GRID = "eth", "12x10" # the p150 target; deliberately not a command-line option +DEFAULT_INPUT = HERE.parent / "samples" / "kashiwanoha_dense.npz" +DEFAULT_EXPECT = "" # unused for a planner (no detections); kept for the shared command line +# Agreement with the stored CPU reference (names of ttaw_smoke.DEFAULT_GATES): the ego trajectory's average / final +# displacement (the predicted_agents array is reported, not gated). Keep them in line with tests/test_e2e_device.py +# (ego mean error <= 0.3 m; the max error <= 1.0 m gate needs the whole array and lives in the device test). +REFERENCE_GATES: Dict[str, Optional[float]] = {"max_ade": 0.3, "max_fde": 1.0} +REQUIRED_KEYS = ("model", "frame_id", "timing_ms", "num_poses", "columns", "trajectory", "turn_indicator", + "predicted_agents") +TRAJECTORY_COLUMNS = ["x", "y", "yaw", "cos", "sin", "velocity", "acceleration"] + + +def build_sample_request(path: Path) -> Dict[str, Any]: + """The ``/predict`` body for the sample: the planner tensors ``.npz`` as ``inputs``.""" + if not Path(path).is_file(): + raise FileNotFoundError(str(path)) + return ttaw_client.build_request(inputs=str(path)) + + +def output_gates(body: Dict[str, Any], expect: str) -> List[str]: + """This model's plausibility gates on a 200 response: 80 trajectory rows of the documented 7 finite columns, + a turn-indicator command in 0..3 with 5 logits, and an encoded predicted-agents array of 80 x 5 per agent.""" + fails: List[str] = [] + traj = body.get("trajectory") or [] + if body.get("columns") != TRAJECTORY_COLUMNS: + fails.append(f"columns {body.get('columns')} != {TRAJECTORY_COLUMNS}") + if len(traj) != 80 or body.get("num_poses") != 80: + fails.append(f"{len(traj)} trajectory rows, expected 80") + if not all(len(row) == 7 and all(math.isfinite(v) for v in row) for row in traj): + fails.append("non-finite or malformed trajectory row") + turn = body.get("turn_indicator") or {} + if turn.get("command") not in (0, 1, 2, 3) or len(turn.get("logits") or []) != 5: + fails.append(f"turn_indicator {turn}") + agents = body.get("predicted_agents") or {} + shape = agents.get("shape") or [] + if len(shape) != 3 or shape[1:] != [80, 5]: + fails.append(f"predicted_agents shape {shape}") + return fails + + +def main(argv: Optional[List[str]] = None) -> int: + ap = argparse.ArgumentParser(description=__doc__.split("\n\n")[0]) + ap.add_argument("--url", default="http://127.0.0.1:20000") + ap.add_argument("--input", type=Path, default=DEFAULT_INPUT, help="sample to POST (default: the shipped one)") + ap.add_argument("--expect", default=DEFAULT_EXPECT, help="unused by this planner (kept for the shared CLI)") + ap.add_argument("--reference", type=Path, help="stored CPU-reference /predict body of --input " + "(default: looked up next to it)") + ap.add_argument("--manifest", type=Path, help="staged tt_kernel_manifest.json: /info must run what it pins") + ap.add_argument("--profile", help="serve profile being smoke-tested (default: the package's default)") + ap.add_argument("--wait", type=float, default=0.0, help="seconds to wait for /health == ok") + ap.add_argument("--out", type=Path, help="write the /predict response here") + a = ap.parse_args(argv) + base = a.url.rstrip("/") + + health = ttaw_client.wait_ready(base, wait_s=a.wait) + if health.get("status") != "ok": + print(ttaw_client.smoke_line(MODEL, f"/health status={health.get('status')!r}", + [f"not ready after {a.wait:.0f} s (error: {health.get('error')})"])) + return 1 + info, fails = ttaw_client.check_service(base, expect_dispatch=EXPECT_DISPATCH, expect_grid=EXPECT_GRID) + if not info: + print(ttaw_client.smoke_line(MODEL, "/info unreachable", fails)) + return 1 + profile = a.profile + if a.manifest: + try: + pinned = ttaw_smoke.pinned_config(a.manifest, a.profile) + except (OSError, ValueError) as e: + fails.append(f"manifest: {e}") + else: + profile = pinned["profile"] + fails += ttaw_smoke.check_pinned(info, pinned, ENV_PREFIX) + + code, body = 0, {} + t0 = time.perf_counter() + try: + code, body = ttaw_client.post(base, build_sample_request(a.input)) + except (OSError, ValueError) as e: # unreadable sample, unknown suffix, server gone + fails.append(f"/predict: {type(e).__name__}: {e}") + rtt_ms = (time.perf_counter() - t0) * 1e3 + if a.out: + a.out.write_text(json.dumps(body, indent=1)) + ref_summary = "none" + if code and code != 200: + fails.append(f"/predict HTTP {code}: {str(body)[:300]}") + elif code == 200: + fails += [f"missing key {k!r}" for k in REQUIRED_KEYS if k not in body] + fails += output_gates(body, a.expect) + ref_path = a.reference or ttaw_smoke.find_reference(a.input, profile, info.get("variant")) + if ref_path: + try: + metrics, more = ttaw_smoke.compare_with_reference(body, ttaw_smoke.load_json(ref_path), + gates=REFERENCE_GATES) + except (OSError, ValueError) as e: + metrics, more = {}, [f"reference {ref_path}: {e}"] + fails += more + ref_summary = f"{ref_path.name} ({ttaw_smoke.describe_metrics(metrics)})" + try: + bad = ttaw_client.check_bad_request(base) + except OSError as e: + bad = f"malformed-request check: {type(e).__name__}: {e}" + if bad: + fails.append(bad) + + device, timing = info.get("device") or {}, body.get("timing_ms") or {} + turn = (body.get("turn_indicator") or {}).get("command_name") + summary = (f"profile={profile or '-'} variant={info.get('variant')} dispatch={device.get('dispatch')} " + f"grid={device.get('grid')} cqs={device.get('num_command_queues')} n={body.get('num_poses')} " + f"turn={turn} " + f"reference={ref_summary} device_ms={timing.get('device')} total_ms={timing.get('total')} " + f"rtt_ms={rtt_ms:.1f}") + print(ttaw_client.smoke_line(MODEL, summary, fails)) + return 1 if fails else 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/code/tt_diffusion_planner/tests/__init__.py b/code/tt_diffusion_planner/tests/__init__.py new file mode 100644 index 0000000000000000000000000000000000000000..9881313609aae2c16a19dae665af4268bd54431d --- /dev/null +++ b/code/tt_diffusion_planner/tests/__init__.py @@ -0,0 +1 @@ +# SPDX-License-Identifier: Apache-2.0 diff --git a/code/tt_diffusion_planner/tests/_research.py b/code/tt_diffusion_planner/tests/_research.py new file mode 100644 index 0000000000000000000000000000000000000000..827d692f54d25b58953f22c9b02334ae61516a2d --- /dev/null +++ b/code/tt_diffusion_planner/tests/_research.py @@ -0,0 +1,57 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Access to the verified research scripts of the porting workspace (``research/diffusion-planner/scripts``), the +independent implementations the host code is unit-tested against. Absent in an installed package or the image: +tests that need them skip.""" +from __future__ import annotations + +import importlib.util +import sys +from pathlib import Path +from typing import Optional + +import numpy as np + +PKG = Path(__file__).resolve().parents[1] +RESEARCH = PKG.parents[3] / "research" / "diffusion-planner" +SCRIPTS = RESEARCH / "scripts" +ORT_GOLDENS = RESEARCH / "ort" +FULL_GOLDENS = RESEARCH / "goldens" +SAMPLES = PKG / "samples" +SMALL_GOLDENS = PKG / "tests" / "goldens" + + +def load_script(name: str): + """Import ``research/diffusion-planner/scripts/.py`` by path (None when absent). Its directory is put on + ``sys.path`` only while it imports (``dp_scene`` imports ``dp_common`` by name).""" + path = SCRIPTS / f"{name}.py" + if not path.is_file(): + return None + key = f"_dp_research_{name}" + if key in sys.modules: + return sys.modules[key] + spec = importlib.util.spec_from_file_location(key, path) + mod = importlib.util.module_from_spec(spec) + sys.path.insert(0, str(SCRIPTS)) + try: + spec.loader.exec_module(mod) + finally: + sys.path.remove(str(SCRIPTS)) + sys.modules[key] = mod + return mod + + +def sample_raw(stem: str) -> dict: + with np.load(SAMPLES / f"{stem}.npz") as z: + return {k: np.array(z[k]) for k in z.files} + + +def research_scene_raw(scene: str) -> Optional[dict]: + path = ORT_GOLDENS / f"golden_{scene}.npz" + if not path.is_file(): + return None + with np.load(path) as z: + return {k[len("raw/"):]: np.array(z[k]) for k in z.files if k.startswith("raw/")} + + +def research_scenes() -> list: + return sorted(p.stem[len("golden_"):] for p in ORT_GOLDENS.glob("golden_*.npz")) diff --git a/code/tt_diffusion_planner/tests/goldens/kashiwanoha_dense.outputs.npz b/code/tt_diffusion_planner/tests/goldens/kashiwanoha_dense.outputs.npz new file mode 100644 index 0000000000000000000000000000000000000000..c1ac6da8661eaecaa68c703d4ba65ba94a9ce9c3 --- /dev/null +++ b/code/tt_diffusion_planner/tests/goldens/kashiwanoha_dense.outputs.npz @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:9cd7cbf5c0f06c1d621411fa671b034a2a593bb681894367f79fe5b6ee4485a6 +size 121935 diff --git a/code/tt_diffusion_planner/tests/goldens/straight_road.outputs.npz b/code/tt_diffusion_planner/tests/goldens/straight_road.outputs.npz new file mode 100644 index 0000000000000000000000000000000000000000..d3f064106c8a5247883193a9540c8b76b511b49d --- /dev/null +++ b/code/tt_diffusion_planner/tests/goldens/straight_road.outputs.npz @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:dd5fcd5a2940eb770caa1fb4ef64064ac547d481ebf4a9b4494acfbf24862d2a +size 33196 diff --git a/code/tt_diffusion_planner/tests/stubs.py b/code/tt_diffusion_planner/tests/stubs.py new file mode 100644 index 0000000000000000000000000000000000000000..9a06b911f1e38031e0cdc8d6bee50111ed44fc0d --- /dev/null +++ b/code/tt_diffusion_planner/tests/stubs.py @@ -0,0 +1,118 @@ +# SPDX-License-Identifier: Apache-2.0 +"""A stand-in for :class:`tt_diffusion_planner.api.DiffusionPlanner` with no device and no weights, for the host tests +of the server and the smoke test. It quacks like the real model (``validate_params``, ``__call__``, ``info``, +``device_info``, ``close``) and runs the bundle's REAL host code around a fake network: the inputs are checked +against ``INPUT_SCHEMA`` and pre-processed by ``host.prepare`` (normalization from the v5.0 statistics embedded below), +the "network" extrapolates every agent's current state at a constant normalised speed, and ``host.make_output`` +post-processes it, so the output depends on the input deterministically and exercises the real post-processing.""" +from __future__ import annotations + +from typing import Any, Dict, Optional + +import numpy as np + +from tt_diffusion_planner import io as tio +from tt_diffusion_planner.api import DiffusionPlanner, Output +from tt_diffusion_planner.host import pipeline as hp +from tt_diffusion_planner.reference import config as C +from tt_diffusion_planner.reference.weights import Normalization + +ETH_P150 = {"dispatch": "eth", "dispatch_requested": "eth", "fallback": None, "grid": "12x10", "grid_x": 12, + "grid_y": 10, "cores": 120, "arch": "blackhole", "eth_patch": True, "device_id": 0, + "num_command_queues": DiffusionPlanner.DEVICE_DEFAULTS["num_command_queues"]} + + +def v5_normalization() -> Normalization: + """The v5.0 ``observation_normalizer`` / ``state_normalizer`` (HF diffusion_planner.param.json; SPEC 3.7), so the + stub needs no weights download.""" + def mv(mean, std): + return np.asarray(mean, np.float32), np.asarray(std, np.float32) + + lane_std = [20.0] * 8 + [1.0] * 25 + obs = {"ego_agent_past": mv([10, 0, 0, 0], [20, 20, 1, 1]), + "ego_current_state": mv([10] + [0] * 9, [20, 20, 1, 1, 20, 20, 20, 20, 1, 1]), + "neighbor_agents_past": mv([10] + [0] * 10, [20, 20, 1, 1, 20, 20, 20, 20, 1, 1, 1]), + "static_objects": mv([10] + [0] * 9, [20, 20, 1, 1, 20, 20, 1, 1, 1, 1]), + "lanes": mv([10] + [0] * 32, lane_std), "route_lanes": mv([10] + [0] * 32, lane_std), + "lanes_speed_limit": mv([0], [20]), "route_lanes_speed_limit": mv([0], [20]), + "polygons": mv([10, 0, 0], [20, 20, 1]), "line_strings": mv([10, 0, 0, 0], [20, 20, 1, 1]), + "goal_pose": mv([10, 0, 0, 0], [20, 20, 1, 1])} + mean = np.tile(np.float32([10, 0, 0, 0]), (C.MAX_NUM_AGENTS, 1, 1)) + std = np.tile(np.float32([20, 20, 1, 1]), (C.MAX_NUM_AGENTS, 1, 1)) + return Normalization(obs, mean, std, 5) + + +class StubModel: + """``StubModel(cfg)`` is a valid ``app.state.ttaw.model_factory`` (``cfg`` = ``config_from_env`` output).""" + + device_info: Dict[str, Any] = ETH_P150 + validate_params = DiffusionPlanner.validate_params + normalization = v5_normalization() + + def __init__(self, cfg: Optional[Dict[str, Any]] = None): + self.cfg = dict(cfg or {}) + self.closed = False + + @property + def info(self) -> Dict[str, Any]: + cls = DiffusionPlanner + return {"model": cls.MODEL_NAME, "variant": self.cfg.get("variant") or cls.DEFAULT_VARIANT, + "weights": {"repo": self.cfg.get("model_id") or cls.DEFAULT_REPO, + "revision": self.cfg.get("revision") or cls.DEFAULT_REVISION}, + "device": self.device_info, "input_kind": cls.INPUT_KIND, "point_fields": list(cls.POINT_FIELDS), + "camera_order": list(cls.CAMERA_ORDER), "labels": list(cls.LABELS), + "runtime_params": {k: v[3] for k, v in cls.RUNTIME_PARAMS.items()}, + "input_schema": {k: {"shape": list(s), "dtype": "float32"} for k, (s, _) in cls.INPUT_SCHEMA.items()}} + + @staticmethod + def fake_network(prep: hp.Prepared) -> Dict[str, np.ndarray]: + """Every agent moves along its current heading at 0.25 normalised units (5 m) per second.""" + cs = prep.decoder.current_states # [321, 4] normalised + t = np.arange(C.OUTPUT_T + 1, dtype=np.float32)[None, :, None] * np.float32(0.025) + x0 = np.repeat(cs[:, None, :], C.OUTPUT_T + 1, axis=1).astype(np.float32) + x0[..., 0:1] += t * cs[:, None, 2:3] + x0[..., 1:2] += t * cs[:, None, 3:4] + logit = np.array([-4.0, 1.0 + float(np.tanh(cs[0, 0])), 0.5, 0.25, 0.0], np.float32) + return {"final_x0": x0, "logit": logit} + + def __call__(self, points: Any = None, *, inputs: Any = None, **params: Any) -> Output: + p = self.validate_params(params) + if points is not None: + raise tio.InputError("this model takes only `inputs` (the planner tensors), not ['points']") + if inputs is None: + raise tio.InputError("this model needs `inputs`: the 15 planner tensors of INPUT_SCHEMA") + arrays = tio.load_inputs(inputs) + prep = hp.prepare(arrays, self.normalization.observation) + raw = self.fake_network(prep) + return hp.make_output(raw["final_x0"], raw["logit"], prep, self.normalization, p, + model=DiffusionPlanner.MODEL_NAME, timing_ms={"device": 1.0}) + + def close(self) -> None: + self.closed = True + + +def sample_inputs(seed: int = 0, n_neighbors: int = 3) -> Dict[str, np.ndarray]: + """A small synthetic planner scene in the raw (pre-normalization) layout of INPUT_SCHEMA.""" + rng = np.random.default_rng(seed) + raw = {k: np.zeros(s, np.float32) for k, s in C.INPUT_SHAPES.items()} + t = np.arange(C.INPUT_T + 1, dtype=np.float32) + raw["ego_agent_past"][0, :, 0] = (t - C.INPUT_T) * 0.8 + raw["ego_agent_past"][0, :, 2] = 1.0 + raw["ego_current_state"][0] = (0, 0, 1, 0, 8.0, 0, 0.2, 0, 0, 0) + for i in range(n_neighbors): + x0, y0, v = float(rng.uniform(5, 40)), float(rng.choice([-3.5, 0.0, 3.5])), float(rng.uniform(2, 10)) + nb = raw["neighbor_agents_past"][0, i] + nb[:, 0], nb[:, 1], nb[:, 2] = x0 + (t - C.INPUT_T) * v * 0.1, y0, 1.0 + nb[:, 4], nb[:, 6], nb[:, 7], nb[:, 8] = v, 1.8, 4.5, 1.0 + for k in range(4): + lane = raw["lanes"][0, k] + lane[:, 0], lane[:, 1] = np.linspace(-20 + 20 * k, 20 * k, 20), 0.0 + lane[:19, 2] = lane[1:, 0] - lane[:-1, 0] + lane[:, 5], lane[:, 7], lane[:, 12] = 1.75, -1.75, 1.0 + raw["lanes_speed_limit"][0, k, 0] = 13.9 + raw["route_lanes"][0, :4] = raw["lanes"][0, :4] + raw["route_lanes_speed_limit"][0, :4] = 13.9 + raw["goal_pose"][0] = (150.0, 0.0, 1.0, 0.0) + raw["ego_shape"][0] = (2.79, 4.89, 1.896) + raw["turn_indicators"][:] = 1.0 + return raw diff --git a/code/tt_diffusion_planner/tests/test_api_host.py b/code/tt_diffusion_planner/tests/test_api_host.py new file mode 100644 index 0000000000000000000000000000000000000000..0ad716793c542a8a2baa3de8145322f0b8e155c7 --- /dev/null +++ b/code/tt_diffusion_planner/tests/test_api_host.py @@ -0,0 +1,78 @@ +# SPDX-License-Identifier: Apache-2.0 +"""The Python API's host hooks without a device (no weights needed): ``DiffusionPlanner._prepare`` / ``_postprocess`` +wired to the host code, ``__call__``'s input handling and parameter validation, ``info``, and the build refusing +unknown compile parameters. + + TT_VISIBLE_DEVICES=none python -m pytest -q code/tt_diffusion_planner/tests/test_api_host.py +""" +from __future__ import annotations + +import threading + +import numpy as np +import pytest + +from tt_diffusion_planner import io as tio +from tt_diffusion_planner.api import DiffusionPlanner, Output +from tt_diffusion_planner.reference import config as C +from tt_diffusion_planner.tests.stubs import StubModel, sample_inputs, v5_normalization + + +class _NoDevice(DiffusionPlanner): + """The real hooks with the network replaced by the stub's fake network (no device, no weights).""" + + def _forward(self, prepared): + return StubModel.fake_network(prepared) + + +def _model() -> DiffusionPlanner: + m = _NoDevice.__new__(_NoDevice) + m._lock, m._closed, m._owns_device, m.device = threading.RLock(), False, False, None + m.variant, m.compile_params, m.warm_variants, m.warmup_ms, m.device_info = "default", {}, [], {}, {} + m.normalization = v5_normalization() + return m + + +def test_identity(): + cls = DiffusionPlanner + assert (cls.MODEL_NAME, cls.ENV_PREFIX) == ("diffusion-planner-p150", "DIFFUSION_PLANNER") + assert cls.INPUT_KIND == "planner" + assert cls.DEFAULT_REVISION == "423efde67f5414734da43a7ad856c17ceb8b51aa" and cls.DEFAULT_TAG == "v5.0" + assert set(cls.ALLOW_PATTERNS) == {C.ENCODER_ONNX, C.DECODER_ONNX, C.TURN_INDICATOR_ONNX, C.PARAM_JSON} + assert Output.__name__ == "Trajectory" and cls.LABELS == ("NONE", "DISABLE", "ENABLE_LEFT", "ENABLE_RIGHT", "KEEP") + with pytest.raises(TypeError): + DiffusionPlanner() + + +def test_call_runs_the_host_hooks(): + m = _model() + out = m(inputs=sample_inputs(), stopping_threshold=0.5) + assert isinstance(out, Output) and out.poses.shape == (80, 7) and out.predicted_agents.shape == (3, 80, 5) + assert set(out.timing_ms) >= {"preprocess", "device", "postprocess", "total"} + ref = StubModel()(inputs=sample_inputs(), stopping_threshold=0.5) + np.testing.assert_array_equal(out.poses, ref.poses) # API hooks == the stub's direct host calls + info = m.info + assert info["input_kind"] == "planner" and info["input_schema"]["lanes"]["shape"] == [1, 140, 20, 33] + assert info["dpm_solver_steps"] == 10 and info["input_names"] == list(C.INPUT_NAMES) + + +@pytest.mark.parametrize("kwargs,match", [ + ({}, "needs `inputs`"), + ({"points": np.zeros((4, 4), np.float32), "inputs": sample_inputs()}, "only `inputs`"), + ({"inputs": {k: v for k, v in sample_inputs().items() if k != "lanes"}}, "missing"), + ({"inputs": sample_inputs(), "velocity_smoothing_window": 0}, "outside"), + ({"inputs": sample_inputs(), "return_denoising_steps": 1}, "boolean"), + ({"inputs": sample_inputs(), "score_threshold": 0.3}, "unknown parameter"), +]) +def test_call_refuses_bad_requests(kwargs, match): + with pytest.raises(tio.InputError, match=match): + _model()(**kwargs) + + +def test_build_refuses_unknown_compile_params(): + """``from_pretrained(**compile_params)`` takes only ``precision``: any other name fails before the weights are read + or the graph is built (the numerics options are ``DIFFUSION_PLANNER_*`` knobs, pinned in ``serve.env``).""" + m = DiffusionPlanner.__new__(DiffusionPlanner) + m.weights_path, m.compile_params, m.device = None, {"ln_fp32": "dec.*"}, None + with pytest.raises(TypeError, match=r"unknown compile parameter\(s\) \['ln_fp32'\]"): + DiffusionPlanner._build(m) diff --git a/code/tt_diffusion_planner/tests/test_bundle_host.py b/code/tt_diffusion_planner/tests/test_bundle_host.py new file mode 100644 index 0000000000000000000000000000000000000000..7bc3dd61eaabb5d5d23199767d5af959ddc196fd --- /dev/null +++ b/code/tt_diffusion_planner/tests/test_bundle_host.py @@ -0,0 +1,415 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Host tests of the bundle wiring (no device, no weights): the vendored ttaw copy is intact and every module +imports without ttnn / torch; the identity in api.py / device.py and the numerics knobs of tt/config.py match +tt-model.yaml; the device configuration resolves like the server's; the stdlib client and smoke test work against +this bundle's app served over real HTTP (ETH dispatch and 12x10 asserted, serve-profile pins, agreement with a stored +reference output); the pip project is the repo-root pyproject.toml; code/scripts/container_smoke.sh (run against fake +tt-model / docker tools) keeps its evidence and always stops the container. + + TT_VISIBLE_DEVICES=none python -m pytest -q code/tt_diffusion_planner/tests/test_bundle_host.py +""" +from __future__ import annotations + +import contextlib +import fnmatch +import hashlib +import json +import os +import shutil +import socket +import subprocess +import sys +import threading +from pathlib import Path +from typing import Iterator + +import numpy as np +import pytest + +import tt_diffusion_planner +from tt_diffusion_planner import device as tdev +from tt_diffusion_planner.api import DiffusionPlanner +from tt_diffusion_planner.tests.stubs import ETH_P150, StubModel, sample_inputs + +PKG_DIR = Path(tt_diffusion_planner.__file__).resolve().parent +CODE_DIR, BUNDLE_DIR, TTAW_DIR = PKG_DIR.parent, PKG_DIR.parents[1], PKG_DIR / "ttaw" +PATCH_MARKER = "single_chip_arch_1cq_no_dispatch_s" # added to topology.cpp by patches/tt-metal-eth-dispatch.patch +VENDOR_IGNORE = ("__pycache__", "*.pyc", "*.pyo", ".pytest_cache", "*.bin", "*.log", ".DS_Store", "VENDORED.json") + +BLOCKER = """ +import importlib, importlib.abc, sys +BLOCKED = set(sys.argv[2].split(",")) +class Block(importlib.abc.MetaPathFinder): + def find_spec(self, name, path, target=None): + if name.split(".")[0] in BLOCKED: + raise ImportError("blocked import of " + name) +sys.meta_path.insert(0, Block()) +sys.path.insert(0, sys.argv[1]) +for target in sys.argv[3:]: + module, _, attr = target.partition(":") + mod = importlib.import_module(module) + if attr: + getattr(mod, attr) +leaked = sorted(k for k in sys.modules if k.split(".")[0] in BLOCKED) +print("ok" if not leaked else "leaked: " + " ".join(leaked)) +""" + + +# ------------------------------------------------------------------------------------------ vendored ttaw + +def test_vendored_ttaw_matches_its_manifest(): + """The copy under code/tt_diffusion_planner/ttaw is exactly what common/tools/vendor.py wrote (never edit it + here).""" + manifest = json.loads((TTAW_DIR / "VENDORED.json").read_text()) + recorded = manifest["files"] + actual = {} + for p in sorted(TTAW_DIR.rglob("*")): + rel = p.relative_to(TTAW_DIR) + if p.is_file() and not any(fnmatch.fnmatch(part, pat) for part in rel.parts for pat in VENDOR_IGNORE): + actual[rel.as_posix()] = hashlib.sha256(p.read_bytes()).hexdigest() + problems = ([f"missing: {k}" for k in sorted(set(recorded) - set(actual))] + + [f"extra: {k}" for k in sorted(set(actual) - set(recorded))] + + [f"modified: {k}" for k in sorted(set(recorded) & set(actual)) if recorded[k] != actual[k]]) + assert problems == [], "re-vendor with common/tools/vendor.py instead of editing the copy" + from tt_diffusion_planner import ttaw + + assert manifest["package"] == "ttaw" and manifest["version"] == ttaw.__version__ + + +_P = "tt_diffusion_planner" + + +@pytest.mark.parametrize("blocked,targets", [ + ("ttnn,torch", [f"{_P}:DiffusionPlanner", f"{_P}:Output", f"{_P}:open_device", f"{_P}:load_inputs", + f"{_P}:INPUT_SCHEMA", f"{_P}.host", f"{_P}.host.solver", f"{_P}.reference.config", f"{_P}.device", + f"{_P}.io", f"{_P}.server", f"{_P}.server.smoke_test"]), + # the CPU reference also runs where ttnn is absent (research venv) + ("ttnn", [f"{_P}.reference", f"{_P}.reference.pipeline", f"{_P}.reference.goldens", f"{_P}.reference.ort"]), +]) +def test_imports_have_no_side_effects(blocked, targets): + if blocked == "ttnn,torch": + pytest.importorskip("fastapi") + targets = targets + ["tt_diffusion_planner.server.app:app"] + out = subprocess.run([sys.executable, "-c", BLOCKER, str(CODE_DIR), blocked] + targets, capture_output=True, + text=True) + assert out.returncode == 0 and out.stdout.strip() == "ok", out.stderr[-2000:] + + +# ----------------------------------------------------------------------------------- identity and config + +@pytest.fixture(scope="module") +def manifest(): + yaml = pytest.importorskip("yaml") + path = BUNDLE_DIR / "tt-model.yaml" + if not path.is_file(): + pytest.skip("no tt-model.yaml next to code/ (installed package or container image)") + return yaml.safe_load(path.read_text()) + + +def test_identity_matches_manifest(manifest): + cls = DiffusionPlanner + assert manifest["name"] == cls.MODEL_NAME and manifest["repo"].endswith("/" + cls.MODEL_NAME) + w = manifest["weights"] + assert (w["repo"], w["revision"]) == (cls.DEFAULT_REPO, cls.DEFAULT_REVISION) + assert list(w.get("allow_patterns") or []) == list(cls.ALLOW_PATTERNS or []) + assert manifest["runtime"]["app"] == "tt_diffusion_planner.server.app:app" + env = manifest["serve"]["env"] + assert env["TT_WEIGHTS_REVISION"] == cls.DEFAULT_REVISION + assert int(env["DIFFUSION_PLANNER_NUM_CQS"]) == tdev.DEVICE_DEFAULTS["num_command_queues"] + for profile in [{"name": "serve", "env": {}}] + list(manifest.get("serve_profiles") or []): + merged = {**env, **(profile.get("env") or {})} + assert merged.get("DIFFUSION_PLANNER_DISPATCH") == "eth", f"{profile['name']}: the p150 target is ETH dispatch" + assert merged.get("DIFFUSION_PLANNER_VARIANT") in cls.VARIANTS, profile["name"] + verify = "\n".join(manifest["verify"]) + assert PATCH_MARKER in verify and cls.DEFAULT_REVISION in verify + + +def test_serve_env_pins_the_numerics(manifest): + """The image serves the gated graph: ``serve.env`` pins every numerics knob of ``tt/config.py`` at its published + default (``KNOBS.serve_env()``, the ttaw.knobs convention) and an empty ``DIFFUSION_PLANNER_PRECISION`` (the + ``DEFAULT_PRECISION`` policy, no override rule); no serve profile overrides them (VERIFY_PORT L2).""" + from tt_diffusion_planner.tt.config import KNOBS + + env = manifest["serve"]["env"] + pins = KNOBS.serve_env() + assert set(pins) == {f"DIFFUSION_PLANNER_{k}" for k in + ("LN_FP32", "HIDDEN_FP32", "SPLIT_MATMUL", "ATTN_FP32_ACC", "ATTN_MATMUL")} + assert {k: env.get(k) for k in pins} == pins + assert env.get("DIFFUSION_PLANNER_PRECISION") == "" + for profile in manifest.get("serve_profiles") or []: + assert not set(profile.get("env") or {}) & (set(pins) | {"DIFFUSION_PLANNER_PRECISION"}), profile["name"] + + +def test_device_config_resolution(monkeypatch): + for name in ("DISPATCH", "NUM_CQS", "L1_SMALL", "TRACE_REGION", "WORKER_L1_SIZE"): + monkeypatch.delenv(f"{tdev.ENV_PREFIX}_{name}", raising=False) + monkeypatch.delenv("TT_DEVICE_ID", raising=False) + cfg = tdev.device_config() + assert (cfg.dispatch, cfg.device_id, cfg.allow_fallback) == ("eth", 0, True) + assert {k: getattr(cfg, k) for k in tdev.DEVICE_DEFAULTS} == tdev.DEVICE_DEFAULTS + assert cfg == DiffusionPlanner.device_config() + monkeypatch.setenv("DIFFUSION_PLANNER_DISPATCH", "worker") + monkeypatch.setenv("DIFFUSION_PLANNER_TRACE_REGION", str(1 << 20)) + monkeypatch.setenv("TT_DEVICE_ID", "1") + cfg = tdev.device_config(num_command_queues=2, worker_l1_size=None) + assert (cfg.dispatch, cfg.trace_region_size, cfg.device_id, cfg.num_command_queues) == ("worker", 1 << 20, 1, 2) + assert cfg == DiffusionPlanner.device_config(num_command_queues=2) + assert tdev.device_config(dispatch="eth").dispatch == "eth" # an explicit argument beats the environment + + +# ------------------------------------------------------------------------------- client and smoke test + +def test_client_script_runs_standalone(tmp_path): + """server/client.py runs with any python3 in isolated mode: no package install, no numpy needed.""" + np.savez(tmp_path / "scene.npz", **sample_inputs()) + subprocess.run([sys.executable, "-I", str(PKG_DIR / "server" / "client.py"), + "--inputs", str(tmp_path / "scene.npz"), + "--param", "stopping_threshold=0.5", "--out", str(tmp_path / "req.json")], check=True, + capture_output=True, text=True) + req = json.loads((tmp_path / "req.json").read_text()) + assert req["inputs"]["format"] == "npz" and req["params"] == {"stopping_threshold": 0.5} + + +class WorkerStub(StubModel): + """A server whose ETH open fell back to WORKER dispatch.""" + + device_info = dict(ETH_P150, dispatch="worker", fallback="RuntimeError: no ETH", grid="11x10", grid_x=11, + cores=110) + + +def _free_port() -> int: + with socket.socket() as s: + s.bind(("127.0.0.1", 0)) + return s.getsockname()[1] + + +@contextlib.contextmanager +def _serve(stub_cls) -> Iterator[str]: + """This bundle's real app under uvicorn on a free port with ``stub_cls`` as the model; yields the base URL.""" + uvicorn = pytest.importorskip("uvicorn") + pytest.importorskip("fastapi") + from tt_diffusion_planner.server import app as server + from tt_diffusion_planner.server import smoke_test + + state = server.app.state.ttaw + saved, state.model_factory = state.model_factory, stub_cls + srv = uvicorn.Server(uvicorn.Config(server.app, host="127.0.0.1", port=_free_port(), log_level="warning", + lifespan="on")) + thread = threading.Thread(target=srv.run, daemon=True) + thread.start() + base = f"http://127.0.0.1:{srv.config.port}" + try: + assert smoke_test.ttaw_client.wait_ready(base, wait_s=30, poll_s=0.1)["status"] == "ok" + yield base + finally: + srv.should_exit = True + thread.join(timeout=30) + state.model_factory = saved + + +def _staged_manifest(tmp_path: Path) -> Path: + """A staged tt_kernel_manifest.json (wire format) with a default profile and one whose pins differ.""" + env = {"TT_WEIGHTS_REVISION": DiffusionPlanner.DEFAULT_REVISION, "DIFFUSION_PLANNER_DISPATCH": "eth", + "DIFFUSION_PLANNER_NUM_CQS": str(tdev.DEVICE_DEFAULTS["num_command_queues"]), + "DIFFUSION_PLANNER_VARIANT": DiffusionPlanner.DEFAULT_VARIANT} + wire = {"name": DiffusionPlanner.MODEL_NAME, + "weights": {"repo_id": DiffusionPlanner.DEFAULT_REPO, "revision": DiffusionPlanner.DEFAULT_REVISION}, + "container": {"serve": {"env": env}, "default_profile": None, + "serve_profiles": [{"name": "default", "env": {}}, + {"name": "other", "env": {"DIFFUSION_PLANNER_VARIANT": "not-a-variant"}}]}} + path = tmp_path / "tt_kernel_manifest.json" + path.write_text(json.dumps(wire)) + return path + + +def test_smoke_test_against_live_server(tmp_path, monkeypatch, capsys): + from tt_diffusion_planner.server import smoke_test + + monkeypatch.setenv("TT_MESH_SHAPE", "1x1") + monkeypatch.setenv("TT_MODEL_WEIGHTS_REVISION", DiffusionPlanner.DEFAULT_REVISION) + for name in ("HF_MODEL", "DIFFUSION_PLANNER_VARIANT", "DIFFUSION_PLANNER_DISPATCH", "DIFFUSION_PLANNER_NUM_CQS"): + monkeypatch.delenv(name, raising=False) + sample = tmp_path / "sample.npz" + np.savez(sample, **sample_inputs()) + reference = StubModel()(inputs=str(sample)).to_dict() # the stub plays the CPU reference of the sample + (tmp_path / "sample.reference.json").write_text(json.dumps(reference)) + shifted = json.loads(json.dumps(reference)) + for row in shifted["trajectory"]: + row[0] += 2.0 + (tmp_path / "shifted.json").write_text(json.dumps(shifted)) + args = ["--input", str(sample), "--out", str(tmp_path / "out.json")] + pinned = args + ["--manifest", str(_staged_manifest(tmp_path))] + + with _serve(StubModel) as url: + assert smoke_test.main(["--url", url] + pinned) == 0 + out = capsys.readouterr().out + assert out.startswith("PASS diffusion-planner-p150: profile=default") and "dispatch=eth grid=12x10" in out + assert "reference=sample.reference.json (trajectory ade=0.0000 fde=0.0000" in out + assert json.loads((tmp_path / "out.json").read_text())["num_poses"] == 80 + assert smoke_test.main(["--url", url, "--profile", "other"] + pinned) == 1 + assert "profile 'other' pins DIFFUSION_PLANNER_VARIANT=not-a-variant" in capsys.readouterr().out + assert smoke_test.main(["--url", url, "--reference", str(tmp_path / "shifted.json")] + args) == 1 + assert "trajectory ADE 2.0000 > 0.3" in capsys.readouterr().out + assert smoke_test.main(["--url", url, "--reference", str(tmp_path / "missing.json")] + args) == 1 + assert "reference " in capsys.readouterr().out + assert smoke_test.main(["--url", url, "--input", str(tmp_path / "missing.npz")]) == 1 + assert "/predict: FileNotFoundError" in capsys.readouterr().out # a FAIL line, not a traceback + with _serve(WorkerStub) as url: + assert smoke_test.main(["--url", url] + args) == 1 + out = capsys.readouterr().out + assert out.startswith("FAIL diffusion-planner-p150:") + assert "dispatch is 'worker', expected 'eth'" in out and "grid is '11x10', expected '12x10'" in out + + +# ------------------------------------------------------------------------- pip project and container smoke + +def test_pip_project_at_repo_root(): + """The pip project is the repo-root pyproject.toml over code/ (never code/pyproject.toml: tt-model copies code/ + over the tt-metal tree before building ttnn, PACKAGING_PILOT.md problem 1).""" + tomllib = pytest.importorskip("tomllib") + path = BUNDLE_DIR / "pyproject.toml" + if not (BUNDLE_DIR / "tt-model.yaml").is_file(): + pytest.skip("not a bundle source checkout (installed package or container image)") + assert not (CODE_DIR / "pyproject.toml").exists(), "the pip project lives at the repo root, never in code/" + project = tomllib.loads(path.read_text()) + find = project["tool"]["setuptools"]["packages"]["find"] + assert find["where"] == ["code"] and any(fnmatch.fnmatch("tt_diffusion_planner", pat) for pat in find["include"]) + data = project["tool"]["setuptools"]["package-data"] + assert {"API.md", "VENDORED.json"} <= set(data["tt_diffusion_planner.ttaw"]) + assert {"kernels/*.cpp", "kernels/*.hpp", "kernels/*.h"} <= set(data["tt_diffusion_planner.ttaw.ops"]) + assert (BUNDLE_DIR / project["project"]["readme"]).is_file() + assert project["tool"]["setuptools"]["dynamic"]["version"]["attr"] == "tt_diffusion_planner.__version__" + yaml = pytest.importorskip("yaml") + shipped = [p for e in yaml.safe_load((BUNDLE_DIR / "tt-model.yaml").read_text())["source"]["extra_code"] + for p in e["paths"]] + assert "pyproject.toml" not in shipped and "setup.py" not in shipped + setuptools = pytest.importorskip("setuptools") + found = set(setuptools.find_packages(where=str(CODE_DIR), include=find["include"], exclude=find.get("exclude", []))) + assert {"tt_diffusion_planner", "tt_diffusion_planner.ttaw", "tt_diffusion_planner.ttaw.server", + "tt_diffusion_planner.ttaw.ops", "tt_diffusion_planner.host", "tt_diffusion_planner.reference"} <= found + assert not any(p.startswith("tt_diffusion_planner.tests") for p in found) + + +FAKE_TT_MODEL = r''' +import json, os, signal, subprocess, sys, time, urllib.request +from pathlib import Path +state = Path(os.environ["FAKE_STATE"]) +args = sys.argv[1:] +with open(state / "calls.txt", "a") as f: + f.write(" ".join(args) + "\n") +opt = lambda name: args[args.index(name) + 1] if name in args else None +if args[0] == "serve": + port, manifest = int(opt("--port")), json.load(open(args[-1])) + (state / "container").write_text("tt-model-" + manifest["name"] + "-" + (opt("--profile") or "default")) + (state / "container.log").write_text("boot: Opening device 0 (dispatch eth, 1 CQ)\n") + if os.environ.get("FAKE_SERVE_FAIL"): + sys.exit("fake serve: the boot failed") + srv = subprocess.Popen([sys.executable, str(state / "server.py"), str(port)], start_new_session=True) + (state / "server.pid").write_text(str(srv.pid)) + for _ in range(200): + try: + urllib.request.urlopen("http://127.0.0.1:%d/health" % port, timeout=1) + break + except OSError: + time.sleep(0.05) +elif args[0] == "stop": + if (state / "server.pid").exists(): + os.kill(int((state / "server.pid").read_text()), signal.SIGTERM) + with open(state / "container.log", "a") as f: + f.write("shutdown: traces released, device closed\n") + (state / "stopped").write_text("1") +elif args[0] == "logs": + print("fallback: tt-model logs") +else: + sys.exit(2) +''' + +FAKE_DOCKER = r''' +import os, sys, time +from pathlib import Path +state = Path(os.environ["FAKE_STATE"]) +args = sys.argv[1:] +with open(state / "docker_calls.txt", "a") as f: + f.write(" ".join(args) + "\n") +known = (state / "container").read_text() if (state / "container").exists() else None +if args[0] == "inspect": + sys.exit(0 if args[-1] == known else 1) +if args[0] == "logs" and args[-1] == known: + deadline = time.time() + 20 + while "--follow" in args and not (state / "stopped").exists() and time.time() < deadline: + time.sleep(0.05) + sys.stdout.write((state / "container.log").read_text()) + sys.exit(0) +sys.exit(1) +''' + +FAKE_SERVER = r''' +import json, sys +from http.server import BaseHTTPRequestHandler, HTTPServer +INFO = {"model": "fake", "device": {"dispatch": "eth", "grid": "12x10", "cores": 120}} +class Handler(BaseHTTPRequestHandler): + def _send(self, code, body): + data = json.dumps(body).encode() + self.send_response(code) + self.send_header("Content-Type", "application/json") + self.send_header("Content-Length", str(len(data))) + self.end_headers() + self.wfile.write(data) + def do_GET(self): + self._send(200, {"status": "ok"} if self.path.startswith("/health") else INFO if self.path == "/info" else {}) + def do_POST(self): + self._send(500, {"detail": "fake server"}) + def log_message(self, *args): + pass +HTTPServer(("127.0.0.1", int(sys.argv[1])), Handler).serve_forever() +''' + + +def _fake_tools(tmp_path: Path) -> dict: + """PATH with fake `tt-model` / `docker` / `python3` (the test interpreter) and a stub server: no docker, no chip.""" + state, bin_dir = tmp_path / "state", tmp_path / "bin" + state.mkdir() + bin_dir.mkdir() + (state / "server.py").write_text(FAKE_SERVER) + for name, body in (("tt-model", FAKE_TT_MODEL), ("docker", FAKE_DOCKER)): + (bin_dir / name).write_text(f"#!{sys.executable}\n{body}") + (bin_dir / name).chmod(0o755) + (bin_dir / "python3").write_text(f'#!/bin/sh\nexec "{sys.executable}" "$@"\n') + (bin_dir / "python3").chmod(0o755) + env = {k: v for k, v in os.environ.items() if k not in ("SMOKE_OUT", "HF_TOKEN", "HUGGING_FACE_HUB_TOKEN")} + env.update(PATH=f"{bin_dir}:{env.get('PATH', '/usr/bin:/bin')}", FAKE_STATE=str(state)) + for tool in ("tt-model", "docker", "python3"): # never the real tools: a real serve would claim the chip + assert shutil.which(tool, path=env["PATH"]) == str(bin_dir / tool) + return env + + +@pytest.mark.parametrize("serve_fails", [False, True]) +def test_container_smoke_keeps_evidence_and_stops(tmp_path, serve_fails): + """code/scripts/container_smoke.sh passes --profile through, saves /info once READY and the container log through + the shutdown (whatever the outcome), and always stops the container.""" + staged = tmp_path / "staged" + staged.mkdir() + manifest = _staged_manifest(staged) + env, logs, port = _fake_tools(tmp_path), tmp_path / "evidence", _free_port() + profile = [] if serve_fails else ["--profile", "other"] + if serve_fails: + env["FAKE_SERVE_FAIL"] = "1" + r = subprocess.run(["bash", str(CODE_DIR / "scripts" / "container_smoke.sh"), str(staged), str(port), *profile, + f"--log-dir={logs}"], env=env, capture_output=True, text=True, timeout=180) + assert r.returncode == 1, r.stdout + r.stderr # the stub serves no valid /predict: the smoke FAILS + calls = (tmp_path / "state" / "calls.txt").read_text().splitlines() + assert calls[0] == " ".join(["serve", "--port", str(port), *profile, str(manifest)]) + assert calls[-1] == " ".join(["stop", *profile, str(manifest)]) and len(calls) == 2 + stem = "diffusion-planner-p150" + ("" if serve_fails else "-other") + "-" + files = {p.name.split(".", 1)[1]: p for p in logs.iterdir()} + assert all(p.name.startswith(stem) for p in files.values()), sorted(files) + log = files["container.log"].read_text() + assert "boot: Opening device 0" in log and "shutdown: traces released" in log # followed through the stop + result = json.loads(files["result.json"].read_text()) + assert (result["rc"], result["result"], result["profile"]) == (1, "FAIL", "default" if serve_fails else "other") + if serve_fails: + assert "info.json" not in files + else: + assert json.loads(files["info.json"].read_text())["device"] == {"dispatch": "eth", "grid": "12x10", + "cores": 120} + assert "FAIL diffusion-planner-p150:" in r.stdout diff --git a/code/tt_diffusion_planner/tests/test_e2e_device.gates.json b/code/tt_diffusion_planner/tests/test_e2e_device.gates.json new file mode 100644 index 0000000000000000000000000000000000000000..58d0da3ddb2b4c72a9929344a5cfd936dadace9e --- /dev/null +++ b/code/tt_diffusion_planner/tests/test_e2e_device.gates.json @@ -0,0 +1,34 @@ +{ + "version": 1, + "note": "Frozen at the first green run; never loosen (PLAN.md 0.3).", + "gates": { + "ego.max_err_m": { + "threshold": 1.0, + "direction": "max", + "metric": "max_abs_err", + "value_at_freeze": 0.7464523315429688, + "frozen_at": "2026-10-07T20:19:23+00:00" + }, + "ego.mean_err_m": { + "threshold": 0.3, + "direction": "max", + "metric": "ade", + "value_at_freeze": 0.14303353428840637, + "frozen_at": "2026-10-08T02:28:56+00:00" + }, + "neighbors.median_max_err_m": { + "threshold": 1.5, + "direction": "max", + "metric": "center_err", + "value_at_freeze": 0.08576023578643799, + "frozen_at": "2026-10-08T02:28:56+00:00" + }, + "turn.command_agreement": { + "threshold": 1.0, + "direction": "min", + "metric": "agreement", + "value_at_freeze": 1.0, + "frozen_at": "2026-10-08T02:28:56+00:00" + } + } +} diff --git a/code/tt_diffusion_planner/tests/test_e2e_device.py b/code/tt_diffusion_planner/tests/test_e2e_device.py new file mode 100644 index 0000000000000000000000000000000000000000..50fc965caf2b0afcedf9be105ce47caed5a7f5c8 --- /dev/null +++ b/code/tt_diffusion_planner/tests/test_e2e_device.py @@ -0,0 +1,136 @@ +# SPDX-License-Identifier: Apache-2.0 +"""End-to-end output agreement vs the CPU reference on the shipped and the public sample scenes (needs one p150). + + bin/devrun -t 1800 -- python -m pytest -q -s code/tt_diffusion_planner/tests/test_e2e_device.py + +No dataset-level metric: the gate is agreement with the fp32 CPU reference of the same Autoware network on the same +inputs (PLAN.md 2.12, SPEC 6.3) on every scene with stored reference outputs: the shipped samples and research scenes +(``research/diffusion-planner/goldens/.npz``) and the nuScenes-derived public-data instants +(``research/diffusion-planner/goldens/public//.npz``; inputs in ``public_data/inputs``, CC BY-NC-SA, local +only), all written by ``code/scripts/ref_golden.py``. Without the research directory (an installed package) the two +shipped samples are run through the CPU reference on the fly. Plus API == server bit identity. + +- ego trajectory (post-processed x, y): max displacement <= 1.0 m and mean displacement <= 0.3 m per scene; +- turn-indicator command identical on every scene; +- neighbours (multi-modal futures, SPEC 10.2): the median over valid neighbours of the per-agent max displacement + <= 1.5 m. + +``GATES`` freeze like the PCC gates (``test_e2e_device.gates.json``); keep ``server/smoke_test.py`` +``REFERENCE_GATES`` in line with them. ``DIFFUSION_PLANNER_E2E_REPORT=`` writes the per-scene numbers as JSON. +""" +from __future__ import annotations + +import base64 +import json +import os +from pathlib import Path + +import numpy as np +import pytest + +from tt_diffusion_planner.reference import config as C +from tt_diffusion_planner.ttaw.golden import GateRegistry + +pytestmark = pytest.mark.device + +PKG = Path(__file__).resolve().parents[1] +RESEARCH = PKG.parents[3] / "research" / "diffusion-planner" +GOLDENS = RESEARCH / "goldens" +SAMPLE = PKG / "samples" / "kashiwanoha_dense.npz" +GATES = GateRegistry.for_test(__file__, { + "ego.max_err_m": (1.0, "max", "max_abs_err"), + "ego.mean_err_m": (0.3, "max", "ade"), + "turn.command_agreement": (1.0, "min", "agreement"), + "neighbors.median_max_err_m": (1.5, "max", "center_err"), +}) + + +def _raw(path: Path, prefix: str = ""): + with np.load(path, allow_pickle=False) as z: + return {k: np.array(z[prefix + k]) for k in C.INPUT_NAMES} + + +def _expected(golden: Path): + with np.load(golden, allow_pickle=False) as z: + return z["out.trajectory"], z["out.predicted_agents"], int(z["out.turn_command"]) + + +def _cases(): + """``[(name, raw inputs, (trajectory, predicted_agents, turn_command) or None)]``.""" + cases = [] + for p in sorted((PKG / "samples").glob("*.npz")): + g = GOLDENS / f"{p.stem}.npz" + cases.append((p.stem, _raw(p), _expected(g) if g.is_file() else None)) + for p in sorted((RESEARCH / "ort").glob("golden_*.npz")): + g = GOLDENS / f"{p.stem[len('golden_'):]}.npz" + if g.is_file(): + cases.append((g.stem, _raw(p, "raw/"), _expected(g))) + for g in sorted(GOLDENS.glob("public/*/*.npz")): + src = RESEARCH / "public_data" / "inputs" / g.parent.name / g.name + if src.is_file(): + cases.append((f"{g.parent.name}/{g.stem}", _raw(src, "raw/"), _expected(g))) + return cases + + +@pytest.fixture(scope="module") +def model(device): + from tt_diffusion_planner import DiffusionPlanner + from tt_diffusion_planner.reference.weights import find_weights_dir + + wd = find_weights_dir() # $DIFFUSION_PLANNER_WEIGHTS_DIR > workspace assets > the HF cache (offline) + if wd is None: + pytest.skip("v5.0 weights not found (DIFFUSION_PLANNER_WEIGHTS_DIR, workspace assets or the HF cache)") + with DiffusionPlanner.from_pretrained(device=device, weights_dir=str(wd)) as m: + yield m + + +def test_scene_agreement(model): + ego_max, ego_mean, turn_equal, nb_median = [], [], [], [] + report = {} + reference = None + for name, raw, want in _cases(): + if want is None: # no stored goldens (installed package): the CPU reference of the shipped sample + if reference is None: + from tt_diffusion_planner.reference.pipeline import ReferencePlanner + + reference = ReferencePlanner(threads=4) + out = reference(inputs=raw) + want = (out.poses, out.predicted_agents, out.turn_indicator["command"]) + dev = model(inputs=raw) + d = np.hypot(*(dev.poses[:, :2] - want[0][:, :2]).T) + ego_max.append(float(d.max())) + ego_mean.append(float(d.mean())) + turn_equal.append(dev.turn_indicator["command"] == want[2]) + if want[1].shape[0]: + per_agent = np.hypot(*(dev.predicted_agents[..., :2] - want[1][..., :2]).transpose(2, 0, 1)) + nb_median.append(float(np.median(per_agent.max(axis=1)))) + report[name] = {"ego_max_m": round(ego_max[-1], 4), "ego_mean_m": round(ego_mean[-1], 4), + "turn_equal": bool(turn_equal[-1]), + "nb_median_max_m": round(nb_median[-1], 4) if want[1].shape[0] else None, + "device_ms": round(float(dev.timing_ms.get("device", float("nan"))), 3)} + print(f"{name}: ego max {ego_max[-1]:.3f} m mean {ego_mean[-1]:.3f} m turn {turn_equal[-1]} " + f"nb median-max {nb_median[-1] if want[1].shape[0] else float('nan'):.3f} m") + print(f"{len(ego_max)} scenes") + out = os.environ.get("DIFFUSION_PLANNER_E2E_REPORT") + if out: + Path(out).parent.mkdir(parents=True, exist_ok=True) + Path(out).write_text(json.dumps({"scenes": report, "info": model.info}, indent=1, default=str) + "\n") + GATES.require("ego.max_err_m", max(ego_max)) + GATES.require("ego.mean_err_m", max(ego_mean)) + GATES.require("turn.command_agreement", float(np.mean(turn_equal))) + GATES.require("neighbors.median_max_err_m", max(nb_median)) + + +def test_api_equals_server(model, monkeypatch): + """The server's ``/predict`` handler and ``model(...)`` give identical outputs (same decoders, same graph).""" + pytest.importorskip("fastapi") + from tt_diffusion_planner.server.app import PredictRequest, app + + state = app.state.ttaw + monkeypatch.setattr(state, "model", model) + monkeypatch.setattr(state, "ready", True) + req = PredictRequest(inputs={"format": "npz", "data": base64.b64encode(SAMPLE.read_bytes()).decode()}) + body = state.predict(req) + want = model(inputs=str(SAMPLE)).to_dict() + assert body["trajectory"] == want["trajectory"] and body["turn_indicator"] == want["turn_indicator"] + assert body["predicted_agents"] == want["predicted_agents"] diff --git a/code/tt_diffusion_planner/tests/test_host_postprocess.py b/code/tt_diffusion_planner/tests/test_host_postprocess.py new file mode 100644 index 0000000000000000000000000000000000000000..d374db423c6455449855ceb053f5dbfd1a9422c8 --- /dev/null +++ b/code/tt_diffusion_planner/tests/test_host_postprocess.py @@ -0,0 +1,198 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Host post-processing (no device, no weights): denormalisation, Eigen's quaternion of the unnormalised rotation, +the trajectory velocity / force-stop / acceleration rules, predicted paths, the turn-indicator decision and the +``make_output`` assembly, against the verified research port (``dp_common.py``), hand-made cases of +``postprocessing_utils.cpp`` / ``turn_indicator_manager.cpp``, and the small goldens of the shipped samples. + + TT_VISIBLE_DEVICES=none python -m pytest -q code/tt_diffusion_planner/tests/test_host_postprocess.py +""" +from __future__ import annotations + +import math + +import numpy as np +import pytest + +from tt_diffusion_planner.host import pipeline as hp +from tt_diffusion_planner.host import postprocess as P +from tt_diffusion_planner.reference import config as C +from tt_diffusion_planner.reference.weights import load_param_json +from tt_diffusion_planner.tests import _research as R + +PARAM_JSON = R.RESEARCH.parents[1] / "assets" / "diffusion-planner" / "hf_diffusion_planner" / C.PARAM_JSON +DEFAULTS = {k: spec[3] for k, spec in hp.RUNTIME_PARAMS.items()} + + +@pytest.fixture(scope="module") +def normalization(): + if not PARAM_JSON.is_file(): + pytest.skip("diffusion_planner.param.json not found") + return load_param_json(PARAM_JSON) + + +def _straight(n=80, v=8.0, decel_after=None, dt=0.1): + """A straight ego path at speed v (optionally braking to a stop after ``decel_after`` points).""" + xs, x, speed = [], 0.0, v + for i in range(n): + if decel_after is not None and i >= decel_after: + speed = max(0.0, speed - 1.5) + x += speed * dt + xs.append(x) + p = np.zeros((n, 4), np.float32) + p[:, 0], p[:, 2] = xs, 1.0 + return p + + +# ------------------------------------------------------------------------------------------------- denorm + +def test_denormalize_matches_research_port(normalization): + x = np.random.default_rng(0).normal(size=(321, 81, 4)).astype(np.float32) + mean, std = normalization.state() + out = P.denormalize(x, mean, std) + np.testing.assert_array_equal(out, (x[:, 1:] * std + mean).astype(np.float32)) + np.testing.assert_array_equal(out[5, 7], x[5, 8] * np.float32([20, 20, 1, 1]) + np.float32([10, 0, 0, 0])) + assert P.denormalize(x, mean, std, keep_current_state=True).shape == (321, 81, 4) + dp = R.load_script("dp_common") + if dp is not None: + sm, ss = np.asarray(normalization.state_mean), np.asarray(normalization.state_std) + np.testing.assert_array_equal(out, dp.denormalize(x[None], (sm, ss))[0]) + + +# --------------------------------------------------------------------------------------------- orientation + +def test_quaternion_of_unit_rotations_is_the_half_angle(): + th = np.linspace(-math.pi + 1e-3, math.pi - 1e-3, 721) + q = P.quaternion_from_cos_sin(np.cos(th), np.sin(th)) + np.testing.assert_allclose(np.abs(q[:, 3]), np.abs(np.cos(th / 2)), atol=1e-12) + np.testing.assert_allclose(np.abs(q[:, 2]), np.abs(np.sin(th / 2)), atol=1e-12) + np.testing.assert_allclose(np.linalg.norm(q, axis=1), 1.0, atol=1e-12) + np.testing.assert_allclose(P.tf2_yaw(q), th, atol=1e-12) + assert not q[:, :2].any() + + +def test_quaternion_of_unnormalised_rotations_follows_eigen(): + """Eigen does not normalise: |(c, s)| = r != 1 shifts the heading tf2::getYaw reads (SPEC 4.6.6). Its two + branches (trace = 2c + 1 > 0, else the m22 pivot) agree at the switch c = -0.5 only for unit vectors; for + r != 1 the published heading jumps there, a quirk of the node that is reproduced, not fixed.""" + r, th = 0.9, math.pi / 2 + q = P.quaternion_from_cos_sin(r * math.cos(th), r * math.sin(th)) + assert math.isclose(P.tf2_yaw(q), 2 * math.atan(0.9 / (1 + 0.9 * math.cos(th))), rel_tol=1e-12) + assert abs(P.tf2_yaw(q) - th) > 0.05 # not atan2(sin, cos) + eps = 1e-12 + for s in (math.sqrt(0.75), -math.sqrt(0.75)): # unit vectors: continuous + a, b = P.quaternion_from_cos_sin(-0.5 + eps, s), P.quaternion_from_cos_sin(-0.5 - eps, s) + np.testing.assert_allclose(P.tf2_yaw(a), P.tf2_yaw(b), atol=1e-9) + # r = 0.58 (c = -0.5, s = 0.3): trace branch w = 0.5, z = 0.3 -> atan2(0.3, 0.16); pivot branch + # z = sqrt(3) / 2, w = 0.3 / sqrt(3) -> atan2(0.3, -0.72) + a, b = P.quaternion_from_cos_sin(-0.5 + eps, 0.3), P.quaternion_from_cos_sin(-0.5 - eps, 0.3) + assert math.isclose(P.tf2_yaw(a), math.atan2(0.3, 0.16), rel_tol=1e-9) + assert math.isclose(P.tf2_yaw(b), math.atan2(0.3, -0.72), rel_tol=1e-9) + + +# --------------------------------------------------------------------------------------------- trajectory + +def test_trajectory_constant_speed(): + tr = P.trajectory_from_poses(_straight(v=8.0)) + np.testing.assert_allclose(tr.velocity, 8.0, rtol=0, atol=1e-4) + assert tr.velocity.dtype == np.float32 and tr.acceleration.dtype == np.float32 + np.testing.assert_allclose(tr.acceleration[:-1], 0.0, atol=2e-3) + assert tr.acceleration[-1] == 0 and not tr.force_stop + np.testing.assert_allclose(tr.time_from_start, 0.1 * np.arange(1, 81)) + assert tr.as_columns().shape == (80, 7) + + +def test_trajectory_force_stop_freezes_poses(): + poses = _straight(v=6.0, decel_after=20) + tr = P.trajectory_from_poses(poses, stopping_threshold=0.3) + assert tr.force_stop + stop = int(np.flatnonzero(tr.velocity == 0)[0]) + assert np.all(tr.velocity[stop:] == 0) + np.testing.assert_array_equal(tr.position[stop:], np.broadcast_to(tr.position[stop - 1], tr.position[stop:].shape)) + off = P.trajectory_from_poses(poses, enable_force_stop=False) + assert not off.force_stop and np.all(off.velocity[-7:] == off.velocity[80 - 8]) + with pytest.raises(ValueError, match="velocity_smoothing_window"): + P.trajectory_from_poses(poses, velocity_smoothing_window=80) + + +@pytest.mark.skipif(R.load_script("dp_common") is None, reason="research scripts not present") +@pytest.mark.parametrize("case", ["cruise", "brake", "golden"]) +def test_trajectory_matches_research_port(case): + """``dp_common.trajectory_from_poses`` (float64 velocities, no float rounding before the average): positions after + freezing identical, velocities within float32 rounding.""" + dp = R.load_script("dp_common") + if case == "golden": + p = R.FULL_GOLDENS / "straight_road.npz" + if not p.is_file(): + pytest.skip("full goldens not generated") + with np.load(p) as z: + poses = z["out.trajectory"][:, [0, 1, 3, 4]] + else: + poses = _straight(v=7.0, decel_after=15 if case == "brake" else None) + tr = P.trajectory_from_poses(poses) + xy, vel, acc = dp.trajectory_from_poses(poses[:, :2]) + np.testing.assert_allclose(tr.position[:, :2], xy, rtol=0, atol=1e-12) + np.testing.assert_allclose(tr.velocity, vel, rtol=1e-6, atol=1e-5) + np.testing.assert_allclose(tr.acceleration, acc, rtol=1e-5, atol=1e-3) + + +# ------------------------------------------------------------------------------------------ turn indicator + +@pytest.mark.skipif(R.load_script("dp_common") is None, reason="research scripts not present") +def test_turn_decision_matches_research_port(): + dp = R.load_script("dp_common") + rng = np.random.default_rng(4) + for _ in range(200): + lg = (rng.normal(size=5) * 4).astype(np.float32) + prev = int(rng.integers(1, 4)) + d = P.TurnIndicatorManager().evaluate(lg, 0.0, prev) + cmd, prob = dp.turn_indicator_command(lg, prev_report=prev) + assert d.command == cmd + np.testing.assert_allclose(d.probabilities, prob, rtol=1e-6, atol=1e-7) + + +def test_turn_decision_rules_and_hold(): + m = P.TurnIndicatorManager() + keep = np.array([0, 0, 0, 0, 5.0], np.float32) # KEEP wins even after -1.25 + d = m.evaluate(keep, 1.0, prev_report=3) + assert (d.command, d.keep_selected) == (3, True) # repeats the last report + left = np.array([0, 0, 4.0, 0, 3.0], np.float32) # 4 > 3 - 1.25: ENABLE_LEFT + d = m.evaluate(left, 2.0, prev_report=1) + assert (d.command, d.keep_selected, d.held) == (2, False, False) + assert math.isclose(sum(d.probabilities), 1.0, rel_tol=1e-3) + d = m.evaluate(keep, 2.9, prev_report=1) # within the 1 s hold window + assert (d.command, d.held) == (2, True) + d = m.evaluate(keep, 3.01, prev_report=1) # expired + assert (d.command, d.held) == (1, False) + tie = np.array([1.0, 1.0, 0, 0, 0], np.float32) # std::max_element: the first maximum + assert P.TurnIndicatorManager().evaluate(tie).command == 0 + assert P.TurnIndicatorManager().evaluate(np.zeros(0, np.float32)).command == 1 + + +# ----------------------------------------------------------------------------------------- make_output + +@pytest.mark.parametrize("stem", ["kashiwanoha_dense", "straight_road"]) +def test_make_output_reproduces_small_goldens(normalization, stem): + """From the stored reference final x0 / logits, the post-processing gives exactly the stored outputs.""" + path = R.SMALL_GOLDENS / f"{stem}.outputs.npz" + if not path.is_file(): + pytest.skip("small goldens not generated") + with np.load(path) as z: + g = {k: z[k] for k in z.files} + prep = hp.prepare(R.sample_raw(stem), normalization.observation) + x0 = np.zeros((C.MAX_NUM_AGENTS, C.OUTPUT_T + 1, C.POSE_DIM), np.float32) + x0[g["final_x0.rows"]] = g["final_x0"] + out = hp.make_output(x0, g["logit"], prep, normalization, DEFAULTS, model="diffusion-planner-p150") + np.testing.assert_array_equal(out.poses, g["trajectory"]) + assert out.turn_indicator["command"] == int(g["turn_command"]) + assert out.columns == hp.TRAJECTORY_COLUMNS and out.predicted_agents.shape == (len(prep.neighbor_rows), 80, 5) + body = out.to_dict() + assert body["num_poses"] == 80 and body["frame_id"] == "base_link" + assert body["predicted_agents"]["shape"] == [len(prep.neighbor_rows), 80, 5] + assert body["meta"]["predicted_agent_rows"] == prep.neighbor_rows.tolist() + # params: a shorter window and no force stop change only the velocity columns + alt = hp.make_output(x0, g["logit"], prep, normalization, {**DEFAULTS, "velocity_smoothing_window": 1}) + np.testing.assert_array_equal(alt.poses[:, :5], out.poses[:, :5]) + steps = [x0] * 11 + dbg = hp.make_output(x0, g["logit"], prep, normalization, {**DEFAULTS, "return_denoising_steps": True}, + denoising_steps=steps).to_dict() + assert dbg["meta"]["denoising_steps"]["shape"] == [11, 81, 4] diff --git a/code/tt_diffusion_planner/tests/test_host_preprocess.py b/code/tt_diffusion_planner/tests/test_host_preprocess.py new file mode 100644 index 0000000000000000000000000000000000000000..34b089f2a8289a0b34e42d612f01ff8eaf9e3431 --- /dev/null +++ b/code/tt_diffusion_planner/tests/test_host_preprocess.py @@ -0,0 +1,221 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Host pre-processing (no device, no weights): the node's normalization rule and the encoder's in-graph +pre-processing as host features, against the verified research reference (``dp_common.py``) and against the rules +of ``preprocessing_utils.cpp:34-84`` / SPEC 3.8 on hand-made cases. (The features are proven against the exported +graph itself, tensor by tensor, in ``test_reference_cpu.py``.) + + TT_VISIBLE_DEVICES=none python -m pytest -q code/tt_diffusion_planner/tests/test_host_preprocess.py +""" +from __future__ import annotations + +import json +import math + +import numpy as np +import pytest + +from tt_diffusion_planner.host.features import atan2_onnx, decoder_masks, encoder_features, line_features +from tt_diffusion_planner.host.normalize import FLT_EPSILON, normalize_array, normalize_inputs, speed_masks +from tt_diffusion_planner.host.pipeline import prepare +from tt_diffusion_planner.reference import config as C +from tt_diffusion_planner.reference.weights import load_param_json +from tt_diffusion_planner.tests import _research as R + +PARAM_JSON = None +for _cand in (R.RESEARCH.parents[1] / "assets" / "diffusion-planner" / "hf_diffusion_planner" / C.PARAM_JSON,): + if _cand.is_file(): + PARAM_JSON = _cand + + +@pytest.fixture(scope="module") +def normalization(): + if PARAM_JSON is None: + pytest.skip("diffusion_planner.param.json not found") + return load_param_json(PARAM_JSON) + + +def _zero_raw(): + raw = {k: np.zeros(s, np.float32) for k, s in C.INPUT_SHAPES.items()} + raw["ego_current_state"][0, :4] = (0, 0, 1, 0) + raw["ego_agent_past"][0, :, 2] = 1.0 + raw["turn_indicators"][:] = 1.0 + raw["ego_shape"][0] = (2.79, 4.89, 1.896) + return raw + + +# ------------------------------------------------------------------------------------------- normalization + +def test_normalize_rule_zero_rows_and_columns(): + mean = np.array([10.0, 0.0, 0.0, 0.0], np.float32) + std = np.array([20.0, 20.0, 1.0, 1.0], np.float32) + v = np.array([[0, 0, 0, 0], # padding: untouched + [1e-8, -1e-8, 0, 1e-9], # every |v| < FLT_EPSILON: untouched (kept as is) + [30.0, 2.0, 1.0, 0.0], # a real row + [0, 0, 1.0, 0]], np.float32) # cos = 1 makes the row real: x -> -0.5 + out = normalize_array(v, mean, std) + np.testing.assert_array_equal(out[0], v[0]) + np.testing.assert_array_equal(out[1], v[1]) + np.testing.assert_array_equal(out[2], ((v[2] - mean) / std).astype(np.float32)) + np.testing.assert_array_equal(out[3], np.array([-0.5, 0, 1, 0], np.float32)) + one = normalize_array(np.array([[0.0], [10.0], [-FLT_EPSILON]], np.float32), np.array([0.0]), np.array([20.0])) + np.testing.assert_array_equal(one.ravel(), np.array([0.0, 0.5, -FLT_EPSILON / 20], np.float32)) + with pytest.raises(ValueError, match="Standard deviation is zero"): + normalize_array(v, mean, np.array([20, 20, 0, 1], np.float32)) + with pytest.raises(KeyError, match="Missing key lanes"): + normalize_inputs({"lanes": np.zeros(C.INPUT_SHAPES["lanes"], np.float32)}, {}) + + +def test_normalize_skips_raw_keys(normalization): + raw = _zero_raw() + raw["sampled_trajectories"][0, 0, 5] = (1, 2, 3, 4) + norm = normalize_inputs(raw, normalization.observation) + for k in C.SKIP_NORMALIZATION: + np.testing.assert_array_equal(norm[k], raw[k]) + assert norm["ego_agent_past"][0, 0].tolist() == [-0.5, 0.0, 1.0, 0.0] + + +@pytest.mark.skipif(R.load_script("dp_common") is None, reason="research scripts not present") +@pytest.mark.parametrize("scene", R.research_scenes() or ["none"]) +def test_normalization_and_masks_equal_research_reference(normalization, scene): + dp = R.load_script("dp_common") + raw = R.research_scene_raw(scene) + if raw is None: + pytest.skip("no research scene") + _, obs, _ = dp.load_param_json(str(PARAM_JSON)) + want = dp.normalize_inputs(raw, obs) + got = normalize_inputs(raw, normalization.observation) + for k in want: + np.testing.assert_array_equal(got[k], want[k], err_msg=k) + for k, v in dp.speed_masks(want).items(): + np.testing.assert_array_equal(speed_masks(got)[k], v, err_msg=k) + with np.load(R.ORT_GOLDENS / f"golden_{scene}.npz") as z: # what the ORT goldens were fed + for k in C.ENCODER_INPUTS: + src = got[k] if k in got else speed_masks(got)[k] + np.testing.assert_array_equal(src, z["in/" + k], err_msg=k) + + +# ------------------------------------------------------------------------------------------------- features + +def test_atan2_onnx_quadrants_and_quirks(): + rng = np.random.default_rng(0) + y = rng.normal(size=1000).astype(np.float32) + x = rng.normal(size=1000).astype(np.float32) + np.testing.assert_allclose(atan2_onnx(y, x), np.arctan2(y, x), rtol=0, atol=2e-6) + # x = 0: dy / 0 = +-inf -> +-pi/2 like atan2 + np.testing.assert_allclose(atan2_onnx(np.float32([2.0, -3.0]), np.float32([0.0, 0.0])), + [math.pi / 2, -math.pi / 2], atol=1e-7) + # the exported decomposition maps y = 0, x < 0 to -pi (IEEE atan2 gives +pi); 0 / 0 is NaN (masked later) + assert atan2_onnx(np.float32(0.0), np.float32(-1.0)) == np.float32(-3.1415927) + assert np.isnan(atan2_onnx(np.float32(0.0), np.float32(0.0))) + + +def test_line_features_append_point_deltas(): + x = np.zeros((2, 4, 3), np.float32) + x[0, :, 0] = [0, 1, 3, 6] + x[0, :, 1] = [0, -1, -1, 2] + x[0, :, 2] = 1.0 + f = line_features(x) + assert f.shape == (2, 4, 5) + np.testing.assert_array_equal(f[0, :, 3], [1, 2, 3, 0]) + np.testing.assert_array_equal(f[0, :, 4], [-1, 0, 3, 0]) + np.testing.assert_array_equal(f[1], 0) + + +def test_encoder_features_rules(normalization): + raw = _zero_raw() + raw["ego_agent_past"][0, :, 0] = np.arange(31) - 30.0 # x: -30 .. 0 + nb = raw["neighbor_agents_past"][0] + nb[0, :, :4] = (12.0, 3.0, 1.0, 0.0) # agent 0: valid everywhere + nb[0, :, 4:8] = (5.0, 1.0, 1.8, 4.5) + nb[0, :, 8] = 1.0 + nb[1, 10, 0] = 50.0 # agent 1: only an old (dropped) row + nb[2, 30, 4] = 2.0 # agent 2: only a velocity at t = 30 + nb[3, 27, :4] = (20.0, 1.0, 0.0, 1.0) # agent 3: a pose at t = 27 only + raw["lanes"][0, 0, :, 0] = np.linspace(0, 19, 20) # lane 0 along +x + raw["lanes"][0, 0, :19, 2] = 1.0 + raw["lanes"][0, 0, 0, 12] = 1.0 # no traffic light (dim 12) + raw["lanes_speed_limit"][0, 0, 0] = 8.0 + raw["polygons"][0, 0, :, 0] = np.linspace(0, 39, 40) + raw["polygons"][0, 0, :, 2] = 1.0 + raw["line_strings"][0, 1, :, 1] = np.linspace(5, 24, 20) + raw["line_strings"][0, 1, :, 3] = 1.0 # road border + norm = normalize_inputs(raw, normalization.observation) + norm.update(speed_masks(norm)) + f = encoder_features(norm, norm) + # ego: the 6 OLDEST rows kept, the rest zeroed; the token is valid with an all-zero position + np.testing.assert_array_equal(f.ego[:6], norm["ego_agent_past"][0, :6]) + assert not f.ego[6:].any() and f.valid["ego"].all() and not f.pos[0, :4].any() and f.pos[0, 4] == 1 + # neighbours: rows 0..24 dropped; agent 1 (old row only) invalid; agent 2's t = 30 row is non-zero (a velocity), + # so the normalisation moved its x to (0 - 10) / 20 = -0.5; the velocity is zeroed after the validity test and + # the valid-step flag set + assert f.valid["neighbor"][:4].tolist() == [True, False, True, True] + assert not f.neighbor[:, :25].any() + np.testing.assert_array_equal(f.neighbor[0, 25:, 4:6], 0) + np.testing.assert_array_equal(f.neighbor[0, 25:, 8], 1) + np.testing.assert_array_equal(f.neighbor[2, 30], [-0.5, 0, 0, 0, 0, 0, 0, 0, 1]) + np.testing.assert_array_equal(f.neighbor[3, :, 8], (np.arange(31) == 27).astype(np.float32)) + np.testing.assert_array_equal(f.neighbor_type[0], [1, 0, 0]) + sl = C.TOKEN_SLICES + np.testing.assert_allclose(f.pos[sl["neighbor"].start, :4], [0.1, 0.15, 1.0, 0.0], atol=1e-7) + # lanes: heading of point 10 from (dx, dy); speed and its TRT mask; attributes from point 0 + assert f.valid["lane"][0] and not f.valid["lane"][1:].any() + np.testing.assert_allclose(f.pos[sl["lane"].start, :4], [norm["lanes"][0, 0, 10, 0], 0, 1, 0], atol=1e-7) + assert f.lane_has_speed[0, 0] and not f.lane_has_speed[1:].any() and f.lane_speed[0, 0] == np.float32(0.4) + assert f.lane_attr[0, 12 - C.LANE_FEATURE_DIM] == 1 + # polygons: the pseudo-heading atan2(dx, is_intersection_area) of point 20 (dx = 1/20 normalised) + h = math.atan2(1.0 / 20.0, 1.0) + np.testing.assert_allclose(f.pos[sl["polygon"].start, 2:4], [math.cos(h), math.sin(h)], atol=1e-6) + # line strings: atan2(is_road_border, is_stop_line) = pi/2 for a road border, whatever its geometry + np.testing.assert_allclose(f.pos[sl["line_string"].start + 1, 2:4], [0, 1], atol=1e-6) + assert not f.valid["line_string"][0] and f.valid["line_string"][1] + # static objects are zeros in Autoware: never valid; goal / ego shape / turn always valid + assert not f.valid["static"].any() and f.valid["goal"].all() and f.valid["turn"].all() + np.testing.assert_array_equal(f.pos[sl["turn"].start, :4 + 9], [0, 0, 1, 0] + [0] * 8 + [1]) + np.testing.assert_array_equal(f.turn, raw["turn_indicators"][0, :30]) + # masks: invalid tokens have no position feature; the ego key is always valid + assert not f.pos[~f.token_valid].any() + assert f.key_valid[0] and (f.key_valid[1:] == f.token_valid[1:]).all() + assert f.token_valid.sum() == 1 + 3 + 1 + 1 + 1 + 3 + # the decoder keys on the CURRENT (normalised) pose only (dit.py:155-156): agent 3, whose only pose is at t = 27, + # is an encoder entity but not a decoder key; agent 2's normalised x = -0.5 makes it a key + d = decoder_masks(norm) + assert d.agent_valid[:5].tolist() == [True, True, False, True, False] + np.testing.assert_array_equal(d.current_states[0], [-0.5, 0, 1, 0]) + np.testing.assert_array_equal(d.current_states[1], norm["neighbor_agents_past"][0, 0, 30, :4]) + + +def test_prepare_on_the_shipped_samples(normalization): + for stem, n_agents in (("kashiwanoha_dense", 88), ("straight_road", 12)): + raw = R.sample_raw(stem) + p = prepare(raw, normalization.observation) + assert p.neighbor_rows.tolist() == list(range(n_agents)) + assert p.enable_force_stop and p.prev_report == 1 + assert p.features.valid["neighbor"].sum() == n_agents == p.decoder.agent_valid.sum() - 1 + np.testing.assert_array_equal(p.x_T, raw["sampled_trajectories"][0]) + assert set(p.norm) >= set(C.INPUT_NAMES) | {"lanes_has_speed_limit", "route_lanes_has_speed_limit"} + + +def test_prepare_refuses_bad_inputs(normalization): + from tt_diffusion_planner.ttaw.io import InputError + + raw = _zero_raw() + del raw["delay"] + with pytest.raises(InputError, match="missing"): + prepare(raw, normalization.observation) + raw = _zero_raw() + raw["lanes"] = np.zeros((1, 70, 20, 33), np.float32) + with pytest.raises(InputError, match="shape"): + prepare(raw, normalization.observation) + + +def test_param_json_contract(normalization): + assert normalization.major_version == 5 + mean, std = normalization.state() + assert mean.shape == (C.MAX_NUM_AGENTS, 1, 4) + np.testing.assert_array_equal(np.unique(mean.reshape(-1, 4), axis=0), [[10, 0, 0, 0]]) + np.testing.assert_array_equal(np.unique(std.reshape(-1, 4), axis=0), [[20, 20, 1, 1]]) + assert set(normalization.observation) == set(C.INPUT_NAMES) - set(C.SKIP_NORMALIZATION) + args = normalization.args + assert (args["agent_num"], args["time_len"], args["future_len"], args["hidden_dim"]) == (320, 31, 80, 256) + assert (args["encoder_mixer_depth"], args["encoder_fusion_depth"], args["decoder_depth"]) == (6, 6, 3) + assert json.loads(json.dumps(args))["diffusion_model_type"] == "x_start" diff --git a/code/tt_diffusion_planner/tests/test_host_solver.py b/code/tt_diffusion_planner/tests/test_host_solver.py new file mode 100644 index 0000000000000000000000000000000000000000..2387dd516f09d79ee5d868a59146764518be3624 --- /dev/null +++ b/code/tt_diffusion_planner/tests/test_host_solver.py @@ -0,0 +1,168 @@ +# SPDX-License-Identifier: Apache-2.0 +"""DPM-Solver++(2M) host code (no device, no weights): the schedule of ``dpm_solver.cpp`` bit for bit, the loop +against the verified research port (``dp_common.dpm_solver_sample``) and against an independent float64 transcription +of the update formulas, the prefix constraint, and the coefficient plan the on-device loop will use. + + TT_VISIBLE_DEVICES=none python -m pytest -q code/tt_diffusion_planner/tests/test_host_solver.py +""" +from __future__ import annotations + +import math + +import numpy as np +import pytest + +from tt_diffusion_planner.host import solver as S +from tt_diffusion_planner.reference import config as C +from tt_diffusion_planner.tests import _research as R + + +def test_timesteps_steps10_match_the_spec(): + ts = S.log_snr_timesteps(10) + assert len(ts) == 11 and all(isinstance(t, np.float32) for t in ts) + np.testing.assert_allclose(np.asarray(ts, np.float64), C.DPM_TIMESTEPS_STEPS10, rtol=0, atol=6e-9) + assert ts[0] == np.float32(1.0) and ts[-1] == np.float32(0.0010006181) + assert np.all(np.diff(np.asarray(ts)) < 0) + + +def test_schedule_against_float64_formulas(): + """The float32 schedule follows the exact VP-linear schedule up to float32 rounding; near t = 0 the C++ float + expression ``1 - exp(2 lmc)`` cancels (lmc ~ -5.5e-5 at t = 1e-3), which costs ~1e-4 relative in sigma and is + reproduced on purpose.""" + b0, b1 = 0.1, 20.0 + for t in (1.0, 0.5, 0.05, 0.001): + tol = 1e-5 if t > 0.01 else 3e-4 + lmc = -0.25 * t * t * (b1 - b0) - 0.5 * t * b0 + assert math.isclose(S.marginal_log_mean_coeff(t), lmc, rel_tol=1e-6) + assert math.isclose(S.marginal_alpha(t), math.exp(lmc), rel_tol=1e-6) + assert math.isclose(S.marginal_std(t), math.sqrt(1 - math.exp(2 * lmc)), rel_tol=tol) + lam = lmc - 0.5 * math.log(1 - math.exp(2 * lmc)) + assert math.isclose(S.marginal_lambda(t), lam, rel_tol=tol) + assert math.isclose(S.inverse_lambda(np.float32(lam)), t, rel_tol=10 * tol) + + +@pytest.mark.skipif(R.load_script("dp_common") is None, reason="research scripts not present") +def test_timesteps_equal_research_port(monkeypatch): + """Same operation order as ``dp_common.py``: with numpy's float32 transcendentals (the research port's choice) + every timestep is bit-identical; with the C library (the node's) the deployed steps = 10 still are, other step + counts may move by one ulp.""" + dp = R.load_script("dp_common") + np.testing.assert_array_equal(np.asarray(S.log_snr_timesteps(10)), np.asarray(dp.log_snr_timesteps(10))) + for steps in (2, 5, 20): + np.testing.assert_allclose(np.asarray(S.log_snr_timesteps(steps)), np.asarray(dp.log_snr_timesteps(steps)), + rtol=3e-7, atol=0) + monkeypatch.setattr(S.SCALAR_MATH, "_fns", {}) + for steps in (2, 5, 10, 20): + np.testing.assert_array_equal(np.asarray(S.log_snr_timesteps(steps)), np.asarray(dp.log_snr_timesteps(steps))) + + +def _toy_model(x, t): + """A deterministic x0-predictor stand-in: mixes x with a time-dependent target.""" + target = np.sin(np.arange(x.size, dtype=np.float32).reshape(x.shape) * np.float32(0.01)) + return (np.float32(0.3) * x + np.float32(1.0 - 0.3) * target * np.float32(1.0 + t)).astype(np.float32) + + +def _float64_dpm(x, steps, model, correct): + """Independent float64 transcription of DPM-Solver++(2M) + denoise-to-zero (dpm_solver.cpp:96-234).""" + b0, b1 = 0.1, 20.0 + lmc = lambda t: -0.25 * t * t * (b1 - b0) - 0.5 * t * b0 # noqa: E731 + alpha = lambda t: math.exp(lmc(t)) # noqa: E731 + sigma = lambda t: math.sqrt(1 - math.exp(2 * lmc(t))) # noqa: E731 + lam = lambda t: math.log(alpha(t)) - math.log(sigma(t)) # noqa: E731 + ts = [float(t) for t in S.log_snr_timesteps(steps)] + x = x.astype(np.float64) + correct(x) + ms, tp = [model(x, ts[0])], [1.0] + t = ts[1] + h = lam(t) - lam(tp[-1]) + x = sigma(t) / sigma(tp[-1]) * x - alpha(t) * math.expm1(-h) * ms[-1] + correct(x) + tp.append(t) + ms.append(model(x, t)) + for step in range(2, steps + 1): + t = ts[step] + h0, h = lam(tp[1]) - lam(tp[0]), lam(t) - lam(tp[1]) + phi = math.expm1(-h) + d = (ms[1] - ms[0]) / (h0 / h) + x = sigma(t) / sigma(tp[1]) * x - alpha(t) * phi * ms[1] - 0.5 * alpha(t) * phi * d + correct(x) + tp = [tp[1], t] + if step < steps: + ms = [ms[1], model(x, t)] + x = model(x, 0.001) + correct(x) + return x + + +def test_loop_against_float64_transcription(): + rng = np.random.default_rng(1) + x0 = rng.normal(size=(321, 81, 4)).astype(np.float32) + cs = rng.normal(size=(321, 4)).astype(np.float32) + calls = [] + + def model(x, t): + calls.append(float(t)) + return _toy_model(np.asarray(x, np.float32), float(t)) + + res = S.dpm_solver_sample(x0, model, lambda x: S.apply_prefix_constraint(x, cs), 10) + assert res.nfe == 11 == len(calls) + plan = S.solver_plan(10) + np.testing.assert_allclose(calls, plan.eval_times, rtol=0, atol=0) + assert len(res.denoising_steps) == 11 and res.denoising_timesteps == list(plan.denoising_timesteps) + for x in res.denoising_steps: + np.testing.assert_array_equal(x[:, 0, :], cs) + want = _float64_dpm(x0, 10, lambda x, t: _toy_model(x.astype(np.float32), t).astype(np.float64), + lambda x: S.apply_prefix_constraint(x, cs)) + np.testing.assert_allclose(res.final_x, want, rtol=0, atol=2e-5) + + +@pytest.mark.skipif(R.load_script("dp_common") is None, reason="research scripts not present") +def test_loop_against_research_port(monkeypatch): + """``dp_common.py`` evaluates the scalars with numpy instead of the C library: with the same scalar math the loop + is bit-identical (same update order), with the C library it differs by the scalars' last bits only.""" + dp = R.load_script("dp_common") + rng = np.random.default_rng(2) + x0 = rng.normal(size=(1, 321, 81, 4)).astype(np.float32) + cs = rng.normal(size=(1, 321, 4)).astype(np.float32) + b_final, b_steps, b_ts, b_nfe = dp.dpm_solver_sample(x0, _toy_model, dp.make_prefix_constraint(cs[0]), 10) + a = S.dpm_solver_sample(x0, _toy_model, lambda x: S.apply_prefix_constraint(x, cs), 10) + assert b_nfe == a.nfe == 11 + np.testing.assert_array_equal(np.asarray(a.denoising_timesteps, np.float32), np.asarray(b_ts, np.float32)) + np.testing.assert_allclose(a.final_x, b_final, rtol=0, atol=2e-5) + monkeypatch.setattr(S.SCALAR_MATH, "_fns", {}) + a = S.dpm_solver_sample(x0, _toy_model, lambda x: S.apply_prefix_constraint(x, cs), 10) + np.testing.assert_array_equal(a.final_x, b_final) + for p, q in zip(a.denoising_steps, b_steps): + np.testing.assert_array_equal(p, q) + + +def test_plan_reproduces_the_updates_bit_for_bit(): + """``x = a x - b m0 - c (m0 - m1) / r0`` with the plan's float32 scalars equals first_update / second_update.""" + rng = np.random.default_rng(3) + plan = S.solver_plan(10) + ts = S.log_snr_timesteps(10) + assert [u.order for u in plan.updates] == [1] + [2] * 9 + x, m0, m1 = (rng.normal(size=(321, 81, 4)).astype(np.float32) for _ in range(3)) + u = plan.updates[0] + want = S.first_update(x, m0, C.NOISE_SCHEDULE_T, ts[1]) + got = (np.float32(u.a) * x - np.float32(u.b) * m0).astype(np.float32) + np.testing.assert_array_equal(got, want) + for k in range(2, 11): + u = plan.updates[k - 1] + t1 = C.NOISE_SCHEDULE_T if k == 2 else ts[k - 2] + want = S.second_update(x, (m1, m0), (t1, ts[k - 1]), ts[k]) + d = ((m0 - m1) / np.float32(u.r0)).astype(np.float32) + got = (np.float32(u.a) * x - np.float32(u.b) * m0 - np.float32(u.c) * d).astype(np.float32) + np.testing.assert_array_equal(got, want, err_msg=f"update {k}") + + +def test_scalar_math_is_the_c_library(): + S.marginal_alpha(0.5) + assert S.SCALAR_MATH.source.startswith("libm"), S.SCALAR_MATH.source + + +def test_steps_below_order_raise(): + with pytest.raises(ValueError): + S.solver_plan(1) + with pytest.raises(ValueError): + S.dpm_solver_sample(np.zeros((321, 81, 4), np.float32), _toy_model, lambda x: None, 1) diff --git a/code/tt_diffusion_planner/tests/test_io_host.py b/code/tt_diffusion_planner/tests/test_io_host.py new file mode 100644 index 0000000000000000000000000000000000000000..9935b0cc9a73475f0c551a94f98372512605a567 --- /dev/null +++ b/code/tt_diffusion_planner/tests/test_io_host.py @@ -0,0 +1,104 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Host tests (no device, no weights): the request decoders shared by the API and the server, i.e. the vendored +``ttaw.io`` as this bundle binds it (``io.load_inputs`` / ``io.decode_inputs`` with ``INPUT_SCHEMA``; the full suite of +ttaw.io is common/tests/host/test_io_outputs_host.py). + + TT_VISIBLE_DEVICES=none python -m pytest -q code/tt_diffusion_planner/tests/test_io_host.py +""" +from __future__ import annotations + +import base64 +import io +from pathlib import Path + +import numpy as np +import pytest + +from tt_diffusion_planner import io as tio +from tt_diffusion_planner.api import DiffusionPlanner +from tt_diffusion_planner.reference import config as C +from tt_diffusion_planner.tests.stubs import sample_inputs + +SAMPLES = ("kashiwanoha_dense", "straight_road") + + +def _npz_b64(arrays) -> str: + buf = io.BytesIO() + np.savez_compressed(buf, **arrays) + return base64.b64encode(buf.getvalue()).decode("ascii") + + +def test_schema_is_the_node_input_map(): + """The 15 raw tensors of ``DiffusionPlannerCore::create_input_data`` (core.cpp:414-595), batch 1, float32.""" + assert DiffusionPlanner.INPUT_SCHEMA is C.INPUT_SCHEMA and DiffusionPlanner.INPUT_KIND == "planner" + assert set(C.INPUT_SCHEMA) == { + "sampled_trajectories", "ego_agent_past", "ego_current_state", "neighbor_agents_past", "static_objects", + "lanes", "lanes_speed_limit", "route_lanes", "route_lanes_speed_limit", "polygons", "line_strings", + "goal_pose", "ego_shape", "turn_indicators", "delay"} + assert C.INPUT_SCHEMA["neighbor_agents_past"][0] == (1, 320, 31, 11) + assert C.INPUT_SCHEMA["lanes"][0] == (1, 140, 20, 33) and C.INPUT_SCHEMA["line_strings"][0] == (1, 60, 20, 4) + assert all(np.dtype(dt) == np.float32 for _, dt in C.INPUT_SCHEMA.values()) + assert tio.DEFAULT_POINT_FIELDS == () and DiffusionPlanner.POINT_FIELDS == () + + +@pytest.mark.parametrize("stem", SAMPLES) +def test_shipped_samples_load_in_every_form(stem): + path = Path(__file__).resolve().parents[1] / "samples" / f"{stem}.npz" + arrays = tio.load_inputs(str(path)) + assert set(arrays) == set(C.INPUT_NAMES) and all(a.dtype == np.float32 for a in arrays.values()) + for form in (path.read_bytes(), {k: v.astype(np.float64) for k, v in arrays.items()}, + {"format": "npz", "data": base64.b64encode(path.read_bytes()).decode()}): + again = tio.load_inputs(form) + for k in C.INPUT_NAMES: + np.testing.assert_array_equal(again[k], arrays[k]) + dec = tio.decode_inputs({"format": "npz", "data": _npz_b64(arrays)}) + np.testing.assert_array_equal(dec["lanes"], arrays["lanes"]) + + +def test_json_arrays_form(): + raw = sample_inputs() + spec = {"format": "json", "arrays": {k: v.tolist() for k, v in raw.items()}} + got = tio.decode_inputs(spec) + for k in C.INPUT_NAMES: + np.testing.assert_array_equal(got[k], raw[k]) + + +@pytest.mark.parametrize("mutate,match", [ + (lambda r: r.pop("delay"), "missing"), + (lambda r: r.update(extra=np.zeros(3, np.float32)), "unexpected"), + (lambda r: r.update(lanes=np.zeros((1, 70, 20, 33), np.float32)), "shape"), + (lambda r: r.update(ego_shape=np.zeros(3, np.float32)), "shape"), + (lambda r: r["neighbor_agents_past"].__setitem__((0, 0, 0, 0), np.nan), "non-finite"), + (lambda r: r["goal_pose"].__setitem__((0, 0), np.inf), "non-finite"), +]) +def test_invalid_inputs_raise_input_error(mutate, match): + raw = sample_inputs() + mutate(raw) + with pytest.raises(tio.InputError, match=match): + tio.load_inputs(raw) + with pytest.raises(tio.InputError, match=match): + tio.decode_inputs({"format": "npz", "data": _npz_b64(raw)}) + + +def test_undecodable_envelopes(): + with pytest.raises(tio.InputError): + tio.decode_inputs({"format": "npz", "data": "not base64!"}) + with pytest.raises(tio.InputError): + tio.decode_inputs({"format": "npz", "data": base64.b64encode(b"not an npz").decode()}) + with pytest.raises(tio.InputError, match="format"): + tio.decode_inputs({"format": "pickle", "data": ""}) + with pytest.raises(tio.InputError): + tio.load_inputs(42) + + +@pytest.mark.parametrize("fmt", ["npz", "npz_compressed", "raw", "list"]) +def test_encode_array_roundtrip(fmt): + a = np.arange(2 * 80 * 5, dtype=np.float32).reshape(2, 80, 5) + d = tio.encode_array(a, fmt=fmt, key="predicted_agents") + if fmt.startswith("npz"): + back = np.load(io.BytesIO(base64.b64decode(d["data"])))["predicted_agents"] + elif fmt == "raw": + back = np.frombuffer(base64.b64decode(d["data"]), dtype=d["dtype"]).reshape(d["shape"]) + else: + back = np.asarray(d["data"], dtype=d["dtype"]) + np.testing.assert_array_equal(back, a) diff --git a/code/tt_diffusion_planner/tests/test_pcc_device.gates.json b/code/tt_diffusion_planner/tests/test_pcc_device.gates.json new file mode 100644 index 0000000000000000000000000000000000000000..04f429386e03b9417cd1bc31b90d4bec1aded502 --- /dev/null +++ b/code/tt_diffusion_planner/tests/test_pcc_device.gates.json @@ -0,0 +1,83 @@ +{ + "version": 1, + "note": "Frozen at the first green run; never loosen (PLAN.md 0.3).", + "gates": { + "dec.eval": { + "threshold": 0.999, + "direction": "min", + "metric": "pcc", + "value_at_freeze": 0.9999819924066826, + "frozen_at": "2026-10-07T20:19:12+00:00" + }, + "enc.ego": { + "threshold": 0.999, + "direction": "min", + "metric": "pcc", + "value_at_freeze": 0.9999723269578298, + "frozen_at": "2026-10-07T20:19:12+00:00" + }, + "enc.ego_shape": { + "threshold": 0.999, + "direction": "min", + "metric": "pcc", + "value_at_freeze": 0.9999996839149371, + "frozen_at": "2026-10-07T20:19:12+00:00" + }, + "enc.encoding": { + "threshold": 0.999, + "direction": "min", + "metric": "pcc", + "value_at_freeze": 0.9999374125960734, + "frozen_at": "2026-10-07T20:19:12+00:00" + }, + "enc.goal": { + "threshold": 0.999, + "direction": "min", + "metric": "pcc", + "value_at_freeze": 0.9999906852333513, + "frozen_at": "2026-10-07T20:19:12+00:00" + }, + "enc.lane": { + "threshold": 0.999, + "direction": "min", + "metric": "pcc", + "value_at_freeze": 0.9999923103079454, + "frozen_at": "2026-10-07T20:19:12+00:00" + }, + "enc.line_string": { + "threshold": 0.999, + "direction": "min", + "metric": "pcc", + "value_at_freeze": 0.999997696425274, + "frozen_at": "2026-10-07T20:19:13+00:00" + }, + "enc.neighbor": { + "threshold": 0.999, + "direction": "min", + "metric": "pcc", + "value_at_freeze": 0.9999516402290836, + "frozen_at": "2026-10-07T20:19:13+00:00" + }, + "enc.polygon": { + "threshold": 0.999, + "direction": "min", + "metric": "pcc", + "value_at_freeze": 0.9999940928919739, + "frozen_at": "2026-10-07T20:19:13+00:00" + }, + "enc.route": { + "threshold": 0.999, + "direction": "min", + "metric": "pcc", + "value_at_freeze": 0.9999982672481239, + "frozen_at": "2026-10-07T20:19:13+00:00" + }, + "enc.turn": { + "threshold": 0.999, + "direction": "min", + "metric": "pcc", + "value_at_freeze": 0.9999945512956152, + "frozen_at": "2026-10-07T20:19:13+00:00" + } + } +} diff --git a/code/tt_diffusion_planner/tests/test_pcc_device.py b/code/tt_diffusion_planner/tests/test_pcc_device.py new file mode 100644 index 0000000000000000000000000000000000000000..177d6b7930603234ab40316449153706612e9c6e --- /dev/null +++ b/code/tt_diffusion_planner/tests/test_pcc_device.py @@ -0,0 +1,145 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Per-module PCC gates: each device stage vs the fp32 CPU reference on the same inputs (needs one p150). + + bin/devrun -t 1800 -- python -m pytest -q -s code/tt_diffusion_planner/tests/test_pcc_device.py + +Gates (PLAN.md 2.12, SPEC 6.3): every encoder module on its VALID rows (entities / tokens) PCC >= 0.999, and every +single decoder evaluation, teacher-forced on the reference's own solver inputs and encoding, PCC >= 0.999 (the gate +value is the minimum over the 11 evaluations and the scenes; the t = 0 slot is excluded: the port zeroes it in the +last projection, an exact rewrite because the prefix constraint overwrites it). They are checked on trace REPLAY +outputs of the debug variants of ``tt.model.TtDiffusionPlanner`` (``encoder_taps`` / ``decode_once``), never TT +against TT; ``test_plan_replay_equals_eager`` checks the production ``plan`` trace bit for bit against an eager run. + +``GATES`` (vendored ``ttaw.golden.GateRegistry``) freezes each gate into ``test_pcc_device.gates.json`` next to this +file at its first green run and refuses a looser declaration afterwards (PLAN.md 0.3; ``TTAW_GATES_READONLY=1`` for +verification runs). Changing a frozen gate means editing that JSON by hand and disclosing it in VERIFICATION_.md. + +Goldens: ``research/diffusion-planner/goldens/.npz`` when present (``code/scripts/ref_golden.py``), else the CPU +reference computes them on the fly (``reference.goldens.scene_goldens``). +""" +from __future__ import annotations + +from pathlib import Path + +import numpy as np +import pytest + +from tt_diffusion_planner.reference import config as C +from tt_diffusion_planner.ttaw.golden import GateRegistry +from tt_diffusion_planner.ttaw.metrics import pcc + +pytestmark = pytest.mark.device + +PKG = Path(__file__).resolve().parents[1] +FULL_GOLDENS = PKG.parents[3] / "research" / "diffusion-planner" / "goldens" +SAMPLES = ("kashiwanoha_dense", "straight_road") # every module has valid rows in at least one of them + +# stage -> minimum PCC vs the fp32 reference on valid rows (keep the names stable across rounds) +GATES = GateRegistry.for_test(__file__, { + "enc.ego": 0.999, "enc.neighbor": 0.999, "enc.lane": 0.999, "enc.route": 0.999, "enc.polygon": 0.999, + "enc.line_string": 0.999, "enc.goal": 0.999, "enc.ego_shape": 0.999, "enc.turn": 0.999, + "enc.encoding": 0.999, "dec.eval": 0.999}) +# reported, not gated (diagnostics for the bring-up of each module) +DIAGNOSTICS = (tuple(f"enc.{c}.{p}" for c in ("ego", "neighbor", "lane", "route", "polygon", "line_string") + for p in ("pre", "mixer")) + ("enc.tokens",) + tuple(f"enc.fusion.{i}" for i in range(6))) + + +@pytest.fixture(scope="module") +def weights(): + from tt_diffusion_planner.reference.weights import find_weights_dir, load_weights + + wd = find_weights_dir() + if wd is None: + pytest.skip("weights not found") + return load_weights(wd) + + +@pytest.fixture(scope="module") +def planner(device, weights): + from tt_diffusion_planner.tt.model import TtDiffusionPlanner + + tt = TtDiffusionPlanner(device, weights, debug=True) + tt.capture() + yield tt + tt.release() + + +def _goldens(stem: str): + path = FULL_GOLDENS / f"{stem}.npz" + if path.is_file(): + with np.load(path, allow_pickle=False) as z: + return {k: z[k] for k in z.files if k != "__meta__"} + from tt_diffusion_planner.reference.goldens import scene_goldens + from tt_diffusion_planner.reference.pipeline import ReferencePlanner + + return scene_goldens(ReferencePlanner(threads=4), str(PKG / "samples" / f"{stem}.npz")) + + +@pytest.fixture(scope="module") +def scenes(weights): + from tt_diffusion_planner.host import pipeline as hp + + out = [] + for stem in SAMPLES: + g = _goldens(stem) + raw = {k: g[f"in.{k}"] for k in C.INPUT_NAMES} + out.append((stem, hp.prepare(raw, weights.normalization.observation), g)) + return out + + +@pytest.fixture(scope="module") +def stage_outputs(planner, scenes): + """``{stage: [(device, reference), ...]}`` over the samples, valid rows only (trace replays).""" + out: dict = {} + for stem, prep, g in scenes: + taps = planner.encoder_taps(prep) # trace replay of the encoder debug variant + for name, _ in C.TOKEN_LAYOUT: + rows = np.flatnonzero(g[f"host.valid.{name}"]) + if rows.size: + out.setdefault(f"enc.{name}", []).append((taps[f"enc.{name}"][rows], g[f"enc.{name}"][rows])) + tok = np.flatnonzero(g["host.token_valid"]) + out.setdefault("enc.encoding", []).append((taps["enc.encoding"][tok], g["enc.encoding"][tok])) + out.setdefault("enc.tokens", []).append((taps["enc.tokens"][tok], g["enc.tokens"][tok])) + for i in range(C.FUSION_DEPTH): + out.setdefault(f"enc.fusion.{i}", []).append((taps[f"enc.fusion.{i}"][tok], g[f"enc.fusion.{i}"][tok])) + for c in ("ego", "neighbor", "lane", "route", "polygon", "line_string"): + rows = g[f"enc.{c}.pre.rows"] + for part in ("pre", "mixer"): + if rows.size: + out.setdefault(f"enc.{c}.{part}", []).append((taps[f"enc.{c}.{part}"][rows], g[f"enc.{c}.{part}"])) + arows = g["dec.rows"] + for k, t in enumerate(g["dec.t"]): + got = planner.decode_once(prep, g["dec.x_in"][k], float(t), encoding=g["enc.encoding"]) + out.setdefault("dec.eval", []).append((got[arows][:, 1:], g["dec.out"][k][:, 1:])) + out.setdefault(f"dec.eval.{k}", []).append((got[arows][:, 1:], g["dec.out"][k][:, 1:])) + return out + + +@pytest.mark.parametrize("stage", GATES.names()) +def test_stage_pcc(stage_outputs, stage): + pairs = stage_outputs.get(stage) + assert pairs, f"no valid rows for {stage} in {SAMPLES}" + value = min(pcc(dev, ref) for dev, ref in pairs) + print(GATES.require(stage, value)) + + +@pytest.mark.parametrize("stage", DIAGNOSTICS + tuple(f"dec.eval.{k}" for k in range(C.DPM_SOLVER_STEPS + 1))) +def test_stage_diagnostics(stage_outputs, stage): + for dev, ref in stage_outputs.get(stage, []): + print(stage, f"pcc={pcc(dev, ref):.6f}", f"max_abs={float(np.abs(np.asarray(dev) - ref).max()):.3e}") + + +def test_plan_replay_equals_eager(planner, scenes): + """The production ``plan`` trace: replay == an eager run of the same graph on the same inputs, bit for bit, and + the prefix constraint holds exactly.""" + for stem, prep, g in scenes: + replay = planner.forward(prep) + eager = planner.forward(prep, eager=True) + for key in ("final_x0", "logit"): + np.testing.assert_array_equal(replay[key], eager[key], err_msg=f"{stem}: {key}") + for a, b in zip(replay["denoising_steps"], eager["denoising_steps"]): + np.testing.assert_array_equal(a, b) + np.testing.assert_array_equal(replay["final_x0"][:, 0], prep.decoder.current_states) + rows = g["dec.rows"] + print(stem, f"final_x0 valid-agent pcc={pcc(replay['final_x0'][rows], g['final_x0'][rows]):.6f}", + f"logit max abs={float(np.abs(replay['logit'] - g['turn.logit']).max()):.4f}") diff --git a/code/tt_diffusion_planner/tests/test_reference_cpu.py b/code/tt_diffusion_planner/tests/test_reference_cpu.py new file mode 100644 index 0000000000000000000000000000000000000000..446d11696cca4bf3d50c55c8ef4e6eabd9ed70b0 --- /dev/null +++ b/code/tt_diffusion_planner/tests/test_reference_cpu.py @@ -0,0 +1,292 @@ +# SPDX-License-Identifier: Apache-2.0 +"""The fp32 CPU reference against ONNX Runtime on the deployed v5.0 ONNX files (PLAN.md 0.3 item 1). + + # research venv (torch, onnx, onnxruntime, pytest), from the bundle root: + PYTHONPATH=code OMP_NUM_THREADS=4 tools/research-venv/bin/python -m pytest -q -p no:cacheprovider \ + code/tt_diffusion_planner/tests/test_reference_cpu.py + +Skipped where onnxruntime or the weights are absent. The model has no data-dependent selection (no top-k, no NMS), +so every comparison is a PCC / max-error check; PCC thresholds are >= 0.9999 on the valid rows of every tap: + +- the weight loader re-labels every initializer of the three files exactly once (by consuming node); +- the host features (``host.features``) equal the encoder graph's own pre-processing tensors (masks exactly); +- every encoder module tap (mixer trunks, category outputs, fusion input, 6 fusion blocks, encoding); +- every decoder evaluation, teacher-forced on ORT's own solver inputs and encoding (t-embedding, pre-projection, + 3 DiT blocks, model output), and the turn head on ORT's encoding / final x0; +- end to end (free-running solver on both sides): final x0, logits, turn command, the post-processed ego + trajectory and the predicted neighbour paths; and the stored research goldens of ``dp_reference.py``. +""" +from __future__ import annotations + +from pathlib import Path + +import numpy as np +import pytest + +pytest.importorskip("onnxruntime") +torch = pytest.importorskip("torch") + +from tt_diffusion_planner.host import pipeline as hp # noqa: E402 +from tt_diffusion_planner.reference import config as C # noqa: E402 +from tt_diffusion_planner.reference.ort import OrtPlanner # noqa: E402 +from tt_diffusion_planner.reference.pipeline import ReferencePlanner # noqa: E402 +from tt_diffusion_planner.reference.weights import coverage, find_weights_dir, load_weights # noqa: E402 +from tt_diffusion_planner.ttaw.golden import TapRegistry # noqa: E402 +from tt_diffusion_planner.ttaw.metrics import ade_fde, pcc # noqa: E402 + +WEIGHTS = find_weights_dir() +pytestmark = pytest.mark.skipif(WEIGHTS is None, reason="Diffusion Planner v5.0 weights not found " + "(set DIFFUSION_PLANNER_WEIGHTS_DIR)") + +PKG = Path(__file__).resolve().parents[1] +RESEARCH = PKG.parents[3] / "research" / "diffusion-planner" +PCC_MIN = 0.9999 +THREADS = 4 + + +def _scenes(): + """``{scene: raw inputs}``: the shipped samples, plus the ORT golden scenes of the research directory (when present, + deduplicated by content).""" + out, seen = {}, set() + for p in sorted((PKG / "samples").glob("*.npz")): + with np.load(p) as z: + out[p.stem] = {k: z[k] for k in C.INPUT_NAMES} + for p in sorted((RESEARCH / "ort").glob("golden_*.npz")): + with np.load(p) as z: + raw = {k: z["raw/" + k] for k in C.INPUT_NAMES} + key = b"".join(raw[k].tobytes() for k in C.INPUT_NAMES) + if any(key == b"".join(v[k].tobytes() for k in C.INPUT_NAMES) for v in out.values()): + continue + if key not in seen: + seen.add(key) + out[p.stem[len("golden_"):]] = raw + return out + + +SCENES = _scenes() +MIXERS = {"ego": "ego_encoder", "neighbor": "neighbor_encoder", "lane": "lane_encoder", "route": "route_encoder", + "polygon": "polygon_encoder", "line_string": "line_string_encoder"} +HOST_TAPS = { # host feature -> encoder graph tensor (batch dim dropped where the graph keeps it) + "ego": "/encoder/Concat_output_0", "neighbor": "/encoder/neighbor_encoder/Where_output_0", + "neighbor_type": "/encoder/neighbor_encoder/Reshape_3_output_0", + "static": "/encoder/static_encoder/Where_output_0", + "lane": "/encoder/lane_encoder/Where_2_output_0", "lane_attr": "/encoder/lane_encoder/Reshape_5_output_0", + "lane_speed": "/encoder/lane_encoder/Reshape_3_output_0", + "route": "/encoder/route_encoder/Where_2_output_0", "route_attr": "/encoder/route_encoder/Reshape_5_output_0", + "route_speed": "/encoder/route_encoder/Reshape_3_output_0", + "polygon": "/encoder/polygon_encoder/Where_2_output_0", + "line_string": "/encoder/line_string_encoder/Where_2_output_0", "turn": "/encoder/Slice_4_output_0", +} +MASK_TAPS = {"token_invalid": "/encoder/Concat_4_output_0", "key_invalid": "/encoder/fusion/Concat_output_0", + "pos": "/encoder/Concat_5_output_0"} +ENC_TAPS = {**{f"enc.{c}.pre": f"/encoder/{m}/Transpose_1_output_0" for c, m in MIXERS.items()}, + **{f"enc.{c}.mixer": f"/encoder/{m}/blocks.{C.MIXER_DEPTH - 1}/Add_1_output_0" for c, m in MIXERS.items()}, + "enc.categories": "/encoder/Concat_3_output_0", "enc.tokens": "/encoder/Add_1_output_0", + **{f"enc.fusion.{i}": f"/encoder/fusion/blocks.{i}/Add_1_output_0" for i in range(C.FUSION_DEPTH)}} +DEC_TAPS = {"temb": "/dit/t_embedder/fc2/Add_output_0", "x": "/dit/Add_output_0", + **{f"block{i}": f"/dit/blocks.{i}/Add_8_output_0" for i in range(C.DIT_DEPTH)}, + "agent_invalid": "/dit/Concat_6_output_0"} + + +@pytest.fixture(scope="module") +def weights(): + return load_weights(WEIGHTS) + + +@pytest.fixture(scope="module") +def ref(weights): + return ReferencePlanner(weights=weights, threads=THREADS) + + +@pytest.fixture(scope="module") +def ort_taps(): + return OrtPlanner(WEIGHTS, threads=THREADS, encoder_taps=list(HOST_TAPS.values()) + list(MASK_TAPS.values()) + + list(ENC_TAPS.values()), decoder_taps=list(DEC_TAPS.values()), + turn_taps=["/ReduceMean_output_0"]) + + +@pytest.fixture(scope="module") +def ort_plain(): + return OrtPlanner(WEIGHTS, threads=THREADS) + + +_CACHE: dict = {} + + +def _ref_run(ref, scene): + key = ("ref", scene) + if key not in _CACHE: + taps = TapRegistry() + _CACHE[key] = (ref.run(SCENES[scene], taps=taps, keep_eval_io=True), taps.to_dict()) + return _CACHE[key] + + +def _ort_run(ort_taps, scene): + key = ("ort", scene) + if key not in _CACHE: + _CACHE[key] = ort_taps.run(SCENES[scene], keep_eval_io=True) + return _CACHE[key] + + +def _check_pcc(name, test, ref_arr, rows=None, pcc_min=PCC_MIN): + t, r = np.asarray(test, np.float64), np.asarray(ref_arr, np.float64) + if rows is not None: + t, r = t[rows], r[rows] + if t.size == 0: + return None + v = pcc(t, r) + err = float(np.abs(t - r).max()) + assert v >= pcc_min, f"{name}: PCC {v:.7f} < {pcc_min} (max |err| {err:.3g})" + return v, err + + +# ------------------------------------------------------------------------------------------------- weights + +def test_weights_cover_every_initializer(weights): + cov = coverage(weights) + assert cov["unused"] == [], cov["unused"][:10] + # the only shared tensors are the two biases the export itself deduplicated (SPEC 6.2) + assert cov["duplicates"] == [("encoder.route_encoder.attribute_emb.b", "encoder.route_encoder.speed_limit_emb.b"), + ("encoder.static_encoder.projection.fc1.b", "encoder.static_encoder.projection.fc2.b")] + assert weights.num_parameters() == 14_545_305 # 14,544,921 float initializers + the two deduplicated biases + assert weights.facts["encoder"]["nodes"] == 1530 and weights.facts["decoder"]["nodes"] == 402 + + +# ------------------------------------------------------------------------------------------ host features + +@pytest.mark.parametrize("scene", sorted(SCENES)) +def test_host_features_match_graph(ref, ort_taps, scene): + res, _ = _ref_run(ref, scene) + o = _ort_run(ort_taps, scene) + f = res.prepared.features + for name, tensor in HOST_TAPS.items(): + got = np.asarray(getattr(f, name), np.float32) + want = o.taps[f"enc:{tensor}"].reshape(got.shape) + np.testing.assert_array_equal(got, want, err_msg=f"host feature {name} != {tensor}") + np.testing.assert_array_equal(~f.token_valid, o.taps["enc:/encoder/Concat_4_output_0"].reshape(-1)) + np.testing.assert_array_equal(~f.key_valid, o.taps["enc:/encoder/fusion/Concat_output_0"].reshape(-1)) + pos = o.taps["enc:/encoder/Concat_5_output_0"].reshape(f.pos.shape) + np.testing.assert_allclose(f.pos[f.token_valid], pos[f.token_valid], rtol=0, atol=2e-6) + np.testing.assert_array_equal(~res.prepared.decoder.agent_valid, o.taps["dec0:/dit/Concat_6_output_0"].reshape(-1)) + for k in C.INPUT_NAMES: # normalization: bit-exact with the ORT path's numpy port of preprocessing_utils.cpp + np.testing.assert_array_equal(res.prepared.norm[k], o.norm[k], err_msg=k) + + +# --------------------------------------------------------------------------------------------- encoder + +@pytest.mark.parametrize("scene", sorted(SCENES)) +def test_encoder_taps_match_ort(ref, ort_taps, scene): + res, taps = _ref_run(ref, scene) + o = _ort_run(ort_taps, scene) + f = res.prepared.features + report = {} + for c in MIXERS: + rows = np.flatnonzero(f.valid[c]) + for part in ("pre", "mixer"): + name = f"enc.{c}.{part}" + want = o.taps[f"enc:{ENC_TAPS[name]}"].reshape(taps[name].shape) + report[name] = _check_pcc(name, taps[name], want, rows) + cats = o.taps["enc:/encoder/Concat_3_output_0"][0] + for name, sl in C.TOKEN_SLICES.items(): + rows = np.flatnonzero(f.valid[name]) + report[f"enc.{name}"] = _check_pcc(f"enc.{name}", taps[f"enc.{name}"], cats[sl], rows) + # invalid entities are exactly zero on both sides + inv = np.flatnonzero(~f.valid[name]) + assert np.all(taps[f"enc.{name}"][inv] == 0) and np.all(cats[sl][inv] == 0), name + rows = np.flatnonzero(f.token_valid) + report["enc.tokens"] = _check_pcc("enc.tokens", taps["enc.tokens"], o.taps["enc:/encoder/Add_1_output_0"][0], rows) + for i in range(C.FUSION_DEPTH): + name = f"enc.fusion.{i}" + report[name] = _check_pcc(name, taps[name], o.taps[f"enc:{ENC_TAPS[name]}"][0], rows) + report["enc.encoding"] = _check_pcc("enc.encoding", res.encoding, o.encoding[0], rows) + # padded tokens leave the encoder bit-identical to each other (SPEC 4.6.3): one pad row stands for all + pad = np.flatnonzero(~f.token_valid) + if pad.size > 1: + assert np.array_equal(res.encoding[pad], np.broadcast_to(res.encoding[pad[0]], res.encoding[pad].shape)) + print(scene, {k: (round(v[0], 8), f"{v[1]:.2e}") for k, v in report.items() if v}) + + +# --------------------------------------------------------------------------------------------- decoder + +@pytest.mark.parametrize("scene", sorted(SCENES)) +def test_decoder_evaluations_teacher_forced(ref, ort_taps, scene): + """Each of the 11 decoder calls of ORT's own solver run, replayed through the reference decoder with ORT's + inputs (x, t) and ORT's encoding: the decoder is checked in isolation, at every diffusion time.""" + res, _ = _ref_run(ref, scene) + o = _ort_run(ort_taps, scene) + rows = np.flatnonzero(res.prepared.decoder.agent_valid) + enc = torch.from_numpy(o.encoding[0]) + kv = ref.decoder.cross_kv(enc) + assert len(o.eval_inputs) == C.DPM_SOLVER_STEPS + 1 == len(o.eval_times) + worst = 1.0 + for k, (x, t) in enumerate(zip(o.eval_inputs, o.eval_times)): + taps = TapRegistry() + with torch.no_grad(): + out = ref.decoder.forward(x, t, kv, res.prepared.decoder.agent_valid, taps, prefix="d").numpy() + for name in ("temb", "x", "block0", "block1", "block2"): + v = _check_pcc(f"dec.{k}.{name}", taps[f"d.{name}"], o.taps[f"dec{k}:{DEC_TAPS[name]}"][0], rows) + worst = min(worst, v[0]) + v = _check_pcc(f"dec.{k}.out", out, o.eval_outputs[k], rows) + worst = min(worst, v[0]) + assert float(np.abs(out[rows] - o.eval_outputs[k][rows]).max()) < 1e-3, f"eval {k}" + print(scene, "worst decoder PCC", worst) + + +@pytest.mark.parametrize("scene", sorted(SCENES)) +def test_turn_head_matches_ort(ref, ort_taps, scene): + o = _ort_run(ort_taps, scene) + taps = TapRegistry() + with torch.no_grad(): + logit = ref.turn.forward(torch.from_numpy(o.encoding[0]), o.final_x0[0], taps).numpy() + # float32 mean of 564 rows: the summation order differs (torch vs ORT ReduceMean), ~4e-6 on values ~0.5 + _check_pcc("turn.pool", taps["turn.pool"], o.taps["turn:/ReduceMean_output_0"][0]) + np.testing.assert_allclose(taps["turn.pool"], o.taps["turn:/ReduceMean_output_0"][0], rtol=0, atol=2e-5) + np.testing.assert_allclose(logit, o.logit[0], rtol=0, atol=2e-4) + + +# ------------------------------------------------------------------------------------------ end to end + +@pytest.mark.parametrize("scene", sorted(SCENES)) +def test_end_to_end_matches_ort(ref, ort_plain, scene): + """Free-running reference vs free-running ORT (``ORT_ENABLE_ALL``, as the node): the DPM loop amplifies nothing + beyond float32 noise.""" + res, _ = _ref_run(ref, scene) + o = ort_plain.run(SCENES[scene]) + rows = np.flatnonzero(res.prepared.decoder.agent_valid) + _check_pcc("final_x0", res.final_x0, o.final_x0[0], rows) + assert float(np.abs(res.final_x0[rows] - o.final_x0[0][rows]).max()) < 1e-3 + np.testing.assert_allclose(res.logit, o.logit[0], rtol=0, atol=1e-3) + np.testing.assert_array_equal(np.asarray(res.denoising_timesteps), np.asarray(o.denoising_timesteps)) + params = {k: spec[3] for k, spec in hp.RUNTIME_PARAMS.items()} + a = hp.make_output(res.final_x0, res.logit, res.prepared, ref.normalization, params) + b = hp.make_output(o.final_x0[0], o.logit[0], res.prepared, ref.normalization, params) + assert a.turn_indicator["command"] == b.turn_indicator["command"] + ade, fde = ade_fde(a.poses[:, :2], b.poses[:, :2]) + assert ade < 1e-3 and fde < 5e-3, (ade, fde) + assert np.abs(a.poses - b.poses).max() < 5e-2 # velocity / acceleration are finite differences of positions + if a.predicted_agents.size: + assert np.abs(a.predicted_agents[..., :2] - b.predicted_agents[..., :2]).max() < 1e-2 + + +@pytest.mark.skipif(not (RESEARCH / "ort").is_dir(), reason="research goldens not present") +def test_matches_research_goldens(ref): + """The stored ``dp_reference.py`` goldens (ORT on the shipped ONNX, normalization by ``dp_common.py``).""" + n = 0 + by_content = {b"".join(v[k].tobytes() for k in C.INPUT_NAMES): name for name, v in SCENES.items()} + for p in sorted((RESEARCH / "ort").glob("golden_*.npz")): + z = np.load(p) + raw = {k: z["raw/" + k] for k in C.INPUT_NAMES} + scene = by_content.get(b"".join(raw[k].tobytes() for k in C.INPUT_NAMES)) + res = _ref_run(ref, scene)[0] if scene else ref.run(raw) + for k in C.INPUT_NAMES: + if k in ("delay",): + continue + np.testing.assert_array_equal(res.prepared.norm[k], z["in/" + k], err_msg=f"{p.name}: in/{k}") + rows = np.flatnonzero(res.prepared.decoder.agent_valid) + _check_pcc(f"{p.name}: encoding", res.encoding, z["encoding"][0], + np.flatnonzero(res.prepared.features.token_valid)) + assert float(np.abs(res.final_x0[rows] - z["final_x_normalized"][0][rows]).max()) < 1e-3, p.name + np.testing.assert_allclose(res.logit, z["logit_multi"][0], rtol=0, atol=1e-3, err_msg=p.name) + np.testing.assert_array_equal(np.asarray(res.denoising_timesteps, np.float32), z["denoising_timesteps"]) + n += 1 + assert n >= 1 diff --git a/code/tt_diffusion_planner/tests/test_reference_host.py b/code/tt_diffusion_planner/tests/test_reference_host.py new file mode 100644 index 0000000000000000000000000000000000000000..e0c968b94e1a5dea3699c67cfb7e01c3a82c1fc0 --- /dev/null +++ b/code/tt_diffusion_planner/tests/test_reference_host.py @@ -0,0 +1,176 @@ +# SPDX-License-Identifier: Apache-2.0 +"""The CPU reference without ONNX Runtime (no device; needs torch, onnx and the v5.0 weights, else skipped): + +- it reproduces the small goldens of the shipped samples (stored by ``code/scripts/ref_golden.py``; the reference + itself is proven against ONNX Runtime in ``test_reference_cpu.py``) and the stored ``/predict`` reference bodies; +- the exact rewrites the TT port applies give the as-exported results (``reference.rewrites``): per-step adaLN tables + folded into the LayerNorm affine, hoisted cross K/V, the pad-relative pre-projection island; and the structural + facts they rest on: uniform-time adaLN rows, bit-identical padded encoder tokens, decoder agent buckets. + + TT_VISIBLE_DEVICES=none python -m pytest -q code/tt_diffusion_planner/tests/test_reference_host.py +""" +from __future__ import annotations + +import json + +import numpy as np +import pytest + +torch = pytest.importorskip("torch") +pytest.importorskip("onnx") + +from tt_diffusion_planner.api import DiffusionPlanner # noqa: E402 +from tt_diffusion_planner.host import pipeline as hp # noqa: E402 +from tt_diffusion_planner.host.solver import solver_plan # noqa: E402 +from tt_diffusion_planner.reference import config as C # noqa: E402 +from tt_diffusion_planner.reference import rewrites as RW # noqa: E402 +from tt_diffusion_planner.reference.pipeline import ReferencePlanner # noqa: E402 +from tt_diffusion_planner.reference.weights import coverage, find_weights_dir, load_weights # noqa: E402 +from tt_diffusion_planner.tests import _research as R # noqa: E402 +from tt_diffusion_planner.ttaw.golden import TapRegistry # noqa: E402 +from tt_diffusion_planner.ttaw.metrics import ade_fde # noqa: E402 + +WEIGHTS = find_weights_dir() +pytestmark = pytest.mark.skipif(WEIGHTS is None, reason="Diffusion Planner v5.0 weights not found " + "(set DIFFUSION_PLANNER_WEIGHTS_DIR)") +STEMS = ("kashiwanoha_dense", "straight_road") + + +@pytest.fixture(scope="module") +def ref(): + return ReferencePlanner(weights=load_weights(WEIGHTS), threads=4) + + +_RUNS: dict = {} + + +def _run(ref, stem): + if stem not in _RUNS: + taps = TapRegistry(include=["enc.encoding", "enc.tokens", "dec.0.*"]) + _RUNS[stem] = (ref.run(R.sample_raw(stem), taps=taps, keep_eval_io=True), taps.to_dict()) + return _RUNS[stem] + + +def _output(ref, stem): + """``ReferencePlanner.__call__`` of the sample, from the cached run (the same two host halves).""" + res, _ = _run(ref, stem) + params = {k: spec[3] for k, spec in hp.RUNTIME_PARAMS.items()} + return hp.make_output(res.final_x0, res.logit, res.prepared, ref.normalization, params, + model=DiffusionPlanner.MODEL_NAME) + + +def test_weight_loader_covers_the_export(ref): + cov = coverage(ref.weights) + assert cov["unused"] == [] and len(cov["duplicates"]) == 2 + + +@pytest.mark.parametrize("stem", STEMS) +def test_reproduces_small_goldens(ref, stem): + path = R.SMALL_GOLDENS / f"{stem}.outputs.npz" + if not path.is_file(): + pytest.skip("small goldens not generated (code/scripts/ref_golden.py)") + res, _ = _run(ref, stem) + with np.load(path) as z: + g = {k: z[k] for k in z.files} + rows = g["final_x0.rows"] + np.testing.assert_array_equal(rows, np.flatnonzero(res.prepared.decoder.agent_valid)) + # another torch build may round differently; the solver keeps such float32 noise small + assert np.abs(res.final_x0[rows] - g["final_x0"]).max() < 1e-3 + np.testing.assert_allclose(res.logit, g["logit"], rtol=0, atol=1e-3) + out = _output(ref, stem) + assert out.turn_indicator["command"] == int(g["turn_command"]) + ade, fde = ade_fde(out.poses[:, :2], g["trajectory"][:, :2]) + assert ade < 1e-3 and fde < 5e-3 + + +@pytest.mark.parametrize("stem", STEMS) +def test_stored_reference_bodies_match(ref, stem): + """``samples/.reference.json`` (the smoke test's oracle) is this reference's ``/predict`` body.""" + path = R.SAMPLES / f"{stem}.reference.json" + if not path.is_file(): + pytest.skip("reference body not generated") + stored = json.loads(path.read_text()) + body = _output(ref, stem).to_dict() + assert stored["model"] == body["model"] and stored["columns"] == body["columns"] + assert stored["turn_indicator"]["command"] == body["turn_indicator"]["command"] + a, b = np.asarray(stored["trajectory"]), np.asarray(body["trajectory"]) + assert a.shape == b.shape == (80, 7) and np.abs(a[:, :2] - b[:, :2]).max() < 1e-2 + + +# -------------------------------------------------------------------------------------------- rewrites + +def test_adaln_is_uniform_over_agents_and_folds_into_layernorm(ref): + """With one diffusion time for every agent (multi-step mode) the t-embedding rows are identical, so the + modulation is a per-step constant; the folded decoder equals the exported one at all 11 evaluation times.""" + res, taps = _run(ref, "straight_road") + temb = taps["dec.0.temb"] + assert np.array_equal(temb, np.broadcast_to(temb[0], temb.shape)) + plan = solver_plan(C.DPM_SOLVER_STEPS) + np.testing.assert_allclose(res.eval_times, plan.eval_times, rtol=0, atol=0) + tables = RW.adaln_tables(ref.weights.params, plan.eval_times) + np.testing.assert_allclose(tables.temb[0], temb[0], rtol=0, atol=2e-5) + kv = ref.decoder.cross_kv(torch.from_numpy(res.encoding)) + rows = np.flatnonzero(res.prepared.decoder.agent_valid) + for k in range(len(plan.eval_times)): + with torch.no_grad(): + want = ref.decoder.forward(res.eval_inputs[k], plan.eval_times[k], kv, res.prepared.decoder.agent_valid) + got = RW.decoder_forward_folded(ref.decoder, tables, k, res.eval_inputs[k], kv, + res.prepared.decoder.agent_valid) + err = float((got - want)[rows].abs().max()) + assert err < 2e-5, f"evaluation {k}: folded decoder differs by {err}" + + +def test_cross_kv_hoisting_is_exact(ref): + res, _ = _run(ref, "kashiwanoha_dense") + hoisted = RW.cross_kv(ref.weights.params, res.encoding) + for i, kv in enumerate(ref.decoder.cross_kv(torch.from_numpy(res.encoding))): + np.testing.assert_allclose(hoisted[i], kv.numpy(), rtol=0, atol=1e-5) + + +@pytest.mark.parametrize("category", ["neighbor", "ego"]) +def test_pad_relative_island_is_exact(ref, category): + res, _ = _run(ref, "kashiwanoha_dense") + f = res.prepared.features + x = f.neighbor if category == "neighbor" else f.ego[None] + cst = RW.island_constants(ref.weights.params, category) + t1e, t2e = RW.island_forward_exported(ref.weights.params, category, x) + t1p, t2p = RW.island_forward_pad_relative(ref.weights.params, cst, x) + assert float((t1e - t1p).abs().max()) < 1e-5 and float((t2e - t2p).abs().max()) < 1e-5 + if category == "neighbor": # a padded agent IS the pad agent + pad = int(np.flatnonzero(~f.valid["neighbor"])[0]) + np.testing.assert_allclose(t2e[pad].numpy(), cst.t2_pad, rtol=0, atol=1e-6) + # the offset dominates the valid agents' fc1 output (SPEC 4.6.7: ~97 % of the mean magnitude on straight, + # 89 % on this denser scene) + valid = np.flatnonzero(f.valid["neighbor"]) + share = float(np.abs(cst.t1_pad).mean() / t1e[valid].abs().mean()) + assert share > 0.8 + + +def test_padded_tokens_are_identical(ref): + """Every invalid entity enters the fusion as the zero vector and is masked as a key, so all padded rows of the + encoding are bit-identical (SPEC 4.6.3: the basis of exact compaction with a pad multiplicity).""" + res, taps = _run(ref, "kashiwanoha_dense") + pad = np.flatnonzero(~res.prepared.features.token_valid) + assert pad.size > 100 + assert not taps["enc.tokens"][pad].any() + enc = res.encoding[pad] + assert np.array_equal(enc, np.broadcast_to(enc[0], enc.shape)) + # turn-head mean pool = (sum of valid rows + m_pad * pad row) / 564 + valid = np.flatnonzero(res.prepared.features.token_valid) + pooled = ((res.encoding[valid].astype(np.float64).sum(0) + pad.size * enc[0].astype(np.float64)) + / C.ENCODING_TOKEN_NUM) + np.testing.assert_allclose(pooled, res.encoding.astype(np.float64).mean(0), rtol=0, atol=1e-6) + + +def test_decoder_agent_buckets_are_exact(ref): + """The decoder on the first K >= 1 + valid neighbours agent rows gives the same outputs for those agents + (padded agents are masked keys and nothing else mixes agents; SPEC 4.6.4).""" + res, _ = _run(ref, "straight_road") + nv = int(res.prepared.decoder.agent_valid.sum()) + kv = ref.decoder.cross_kv(torch.from_numpy(res.encoding)) + x, t = res.eval_inputs[3], res.eval_times[3] + with torch.no_grad(): + full = ref.decoder.forward(x, t, kv, res.prepared.decoder.agent_valid).numpy() + for k in sorted({nv, 16, 32, 64}): + part = ref.decoder.forward(x[:k], t, kv, res.prepared.decoder.agent_valid[:k]).numpy() + assert np.abs(part[:nv] - full[:nv]).max() < 2e-5, k diff --git a/code/tt_diffusion_planner/tests/test_server_host.py b/code/tt_diffusion_planner/tests/test_server_host.py new file mode 100644 index 0000000000000000000000000000000000000000..53c67f92c37c9abeb906f827cabe48df8d0e86d3 --- /dev/null +++ b/code/tt_diffusion_planner/tests/test_server_host.py @@ -0,0 +1,116 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Host tests of the HTTP contract (no device, no weights): the real app of this bundle with a stub model that runs +the real host pre- and post-processing around a fake network (``tests/stubs.py``). + + TT_VISIBLE_DEVICES=none python -m pytest -q code/tt_diffusion_planner/tests/test_server_host.py +""" +from __future__ import annotations + +import base64 +import io + +import numpy as np +import pytest + +fastapi = pytest.importorskip("fastapi") +from fastapi.testclient import TestClient # noqa: E402 + +from tt_diffusion_planner.api import DiffusionPlanner # noqa: E402 +from tt_diffusion_planner.reference import config as C # noqa: E402 +from tt_diffusion_planner.server import app as server # noqa: E402 +from tt_diffusion_planner.tests.stubs import StubModel, sample_inputs # noqa: E402 + + +@pytest.fixture +def client(monkeypatch): + monkeypatch.setenv("TT_MESH_SHAPE", "1x1") + monkeypatch.setenv("TT_MODEL_WEIGHTS_REVISION", "0" * 40) + monkeypatch.setattr(server.app.state.ttaw, "model_factory", StubModel) + with TestClient(server.app) as c: + yield c + + +def _npz_b64(arrays) -> str: + buf = io.BytesIO() + np.savez_compressed(buf, **arrays) + return base64.b64encode(buf.getvalue()).decode() + + +def _req(arrays=None, **extra): + return {"inputs": {"format": "npz", "data": _npz_b64(sample_inputs() if arrays is None else arrays)}, **extra} + + +def test_health_info_models(client): + assert client.get("/health").json()["status"] == "ok" + assert client.get("/v1/health").json()["status"] == "ok" + info = client.get("/info").json() + assert info["model"] == DiffusionPlanner.MODEL_NAME and info["weights"]["revision"] == "0" * 40 + assert info["device"]["dispatch"] == "eth" and info["device"]["grid"] == "12x10" + assert info["autoware"]["package"] == "autoware_diffusion_planner" + assert info["output"]["labels"] == list(DiffusionPlanner.LABELS) + assert info["variant"] == DiffusionPlanner.DEFAULT_VARIANT + assert client.get("/v1/models").json()["data"][0]["id"] == "AutowareFoundation/diffusion_planner" + + +def test_predict_ok_and_equals_api(client): + r = client.post("/predict", json=_req(params={"stopping_threshold": 0.4})) + assert r.status_code == 200, r.text + body = r.json() + assert body["model"] == DiffusionPlanner.MODEL_NAME and body["frame_id"] == "base_link" + assert body["num_poses"] == 80 and len(body["trajectory"]) == 80 and len(body["trajectory"][0]) == 7 + assert body["columns"] == ["x", "y", "yaw", "cos", "sin", "velocity", "acceleration"] + assert body["turn_indicator"]["command"] in (0, 1, 2, 3) and len(body["turn_indicator"]["logits"]) == 5 + assert body["predicted_agents"]["shape"] == [3, 80, 5] + assert {"decode", "model_call", "total", "device"} <= set(body["timing_ms"]) + state = server.app.state.ttaw + req = server.PredictRequest(**_req()) + assert state.predict(req)["trajectory"] == state.model(inputs=sample_inputs()).to_dict()["trajectory"] + + +def test_predict_json_arrays_and_npz_output(client): + arrays = {k: v.tolist() for k, v in sample_inputs().items()} + r = client.post("/predict", json={"inputs": {"format": "json", "arrays": arrays}, "output_format": "npz"}) + assert r.status_code == 200, r.text + assert r.json()["arrays"]["poses"]["shape"] == [80, 7] + + +def _bad(key, value): + raw = sample_inputs() + raw[key] = value + return raw + + +@pytest.mark.parametrize("payload,code", [ + ({"inputs": {"format": "npz", "data": "not base64!"}}, 400), # undecodable + (_req(_bad("lanes", np.zeros((1, 70, 20, 33), np.float32))), 400), # wrong shape + (_req(_bad("goal_pose", np.full((1, 4), np.inf, np.float32))), 400), # non-finite + (_req({k: v for k, v in sample_inputs().items() if k != "delay"}), 400), # missing tensor + (_req(params={"velocity_smoothing_window": 0}), 400), # out of range + (_req(params={"velocity_smoothing_window": 80}), 400), + (_req(params={"return_denoising_steps": "yes"}), 400), # not a boolean + (_req(params={"score_threshold": 0.5}), 400), # unknown knob + ({"points": {"format": "list", "values": [[0, 0, 0, 0]]}}, 400), # wrong input kind + ({}, 400), # no inputs + ({"inputz": {}}, 422), # schema violation +]) +def test_predict_errors(client, payload, code): + r = client.post("/predict", json=payload) + assert r.status_code == code, r.text + + +def test_503_while_starting(client, monkeypatch): + monkeypatch.setattr(server.app.state.ttaw, "ready", False) + assert client.post("/predict", json={}).status_code == 503 + assert client.get("/health").json()["status"] == "starting" + + +def test_input_schema_is_published(): + assert {k: list(v[0]) for k, v in DiffusionPlanner.INPUT_SCHEMA.items()} == { + k: list(s) for k, s in C.INPUT_SHAPES.items()} + + +def test_mesh_shape_parser(): + p = server.parse_mesh_shape + assert p("1x1") == p("(1, 1)") == p("1,1") == p(None) == (1, 1) + with pytest.raises(RuntimeError): + p("two by two") diff --git a/code/tt_diffusion_planner/tests/test_tt_graph_host.py b/code/tt_diffusion_planner/tests/test_tt_graph_host.py new file mode 100644 index 0000000000000000000000000000000000000000..2f349e38dfff84d50c64837dfbf337503179404f --- /dev/null +++ b/code/tt_diffusion_planner/tests/test_tt_graph_host.py @@ -0,0 +1,203 @@ +# SPDX-License-Identifier: Apache-2.0 +"""The whole ttnn graph of ``tt.model.TtDiffusionPlanner`` on the FAKE ttnn (numpy numerics: bf16 tensors hold +bf16-rounded values, fp32 tensors float32), against the fp32 CPU reference's goldens: catches wiring mistakes +(shapes, transposes, token order, masks, solver indices, packing) without a device. The device numbers are gated +in ``test_pcc_device.py`` / ``test_e2e_device.py``. + +The fake op set is ``common/tests/host/fake_ttnn.py`` + ``fake_ttnn_cnn.py`` + ``fake_ttnn_attention.py`` (loaded +by path like ``fake_ttnn_plugin``) plus LayerNorm / mean / subtract / tanh-GELU defined here; skipped when the +workspace's common tree is absent (an installed package). + + TT_VISIBLE_DEVICES=none python -m pytest -q -p fake_ttnn_plugin code/tt_diffusion_planner/tests/test_tt_graph_host.py +""" +from __future__ import annotations + +import importlib.util +import math +import sys +from pathlib import Path + +import numpy as np +import pytest + +from tt_diffusion_planner.host import pipeline as hp +from tt_diffusion_planner.reference import config as C +from tt_diffusion_planner.ttaw.metrics import pcc + +PKG = Path(__file__).resolve().parents[1] +FAKES = PKG.parents[3] / "common" / "tests" / "host" +GOLDENS = PKG.parents[3] / "research" / "diffusion-planner" / "goldens" +SCENE = "kashiwanoha_dense" + + +def _load(name: str): + path = FAKES / f"{name}.py" + if not path.is_file(): + pytest.skip(f"{path} not found (workspace fake ttnn op sets)") + if name in sys.modules: + return sys.modules[name] + spec = importlib.util.spec_from_file_location(name, path) + mod = importlib.util.module_from_spec(spec) + sys.modules[name] = mod + spec.loader.exec_module(mod) + return mod + + +def _install_planner_ops(fake): + """LayerNorm, mean, subtract, rsqrt, scalar add and a linear with "gelu_tanh" on top of the CNN / attention + fakes.""" + names = ("linear", "layer_norm", "mean", "subtract", "rsqrt", "add", "gelu", "GeluVariant") + saved = {n: getattr(fake, n) for n in names if hasattr(fake, n)} + base_linear, base_add = fake.linear, fake.add + variants = type("GeluVariant", (), {"Accurate": "accurate", "FastLut": "fast_lut", "Tanh": "tanh"}) + + def gelu_tanh(x): + return 0.5 * x * (1.0 + np.tanh(math.sqrt(2.0 / math.pi) * (x + 0.044715 * x ** 3))) + + def linear(a, b, *, bias=None, activation=None, dtype=None, compute_kernel_config=None, **kw): + if activation != "gelu_tanh": + return base_linear(a, b, bias=bias, activation=activation, dtype=dtype, + compute_kernel_config=compute_kernel_config, **kw) + inputs = [a, b] + ([bias] if bias is not None else []) + + def fn(x, y, *bb): + z = np.matmul(np.asarray(x, np.float64), np.asarray(y, np.float64)) + return gelu_tanh(z + (bb[0].reshape(-1) if bb else 0.0)) + + return fake._op("linear_gelu_tanh", inputs, fn, dtype or a.dtype, fake.TILE_LAYOUT) + + def layer_norm(x, *, epsilon=1e-12, weight=None, bias=None, compute_kernel_config=None, **kw): + inputs = [x] + [t for t in (weight, bias) if t is not None] + + def fn(a, *gb): + a = np.asarray(a, np.float64) + y = (a - a.mean(-1, keepdims=True)) / np.sqrt(a.var(-1, keepdims=True) + epsilon) + if weight is not None: + y = y * np.asarray(gb[0], np.float64).reshape(-1) + if bias is not None: + y = y + np.asarray(gb[-1], np.float64).reshape(-1) + return y + + return fake._op("layer_norm", inputs, fn, x.dtype, x.layout, (float(epsilon),)) + + def mean(x, dim, keepdim=False, **kw): + return fake._op("mean", [x], lambda a: np.asarray(a, np.float64).mean(axis=dim, keepdims=keepdim), x.dtype, + x.layout, (dim, keepdim)) + + def subtract(a, b, **kw): + return fake._op("subtract", [a, b], lambda x, y: np.asarray(x, np.float64) - y, a.dtype, a.layout) + + def rsqrt(x, *, fast_and_approximate_mode=True, **kw): + return fake._op("rsqrt", [x], lambda a: 1.0 / np.sqrt(np.asarray(a, np.float64)), x.dtype, x.layout) + + def add(a, b, **kw): + if isinstance(b, (int, float)): + value = float(b) + return fake._op("add_scalar", [a], lambda x: np.asarray(x, np.float64) + value, a.dtype, a.layout, + (value,)) + return base_add(a, b, **kw) + + def gelu(x, *, variant=None, fast_and_approximate_mode=False, **kw): + if variant == variants.Tanh: + return fake._op("gelu_tanh", [x], lambda a: gelu_tanh(np.asarray(a, np.float64)), x.dtype, x.layout) + erf = np.vectorize(math.erf) + return fake._op("gelu", [x], lambda a: 0.5 * a * (1.0 + erf(np.asarray(a, np.float64) / math.sqrt(2.0))), + x.dtype, x.layout) + + for name, fn in (("linear", linear), ("layer_norm", layer_norm), ("mean", mean), ("subtract", subtract), + ("rsqrt", rsqrt), ("add", add), ("gelu", gelu), ("GeluVariant", variants)): + setattr(fake, name, fn) + + def restore(): + for name in names: + if name in saved: + setattr(fake, name, saved[name]) + elif hasattr(fake, name): + delattr(fake, name) + return restore + + +@pytest.fixture(scope="module") +def fake_planner(): + import ttnn # the fake (fake_ttnn_plugin) + + if not hasattr(ttnn, "_op"): + pytest.skip("needs the fake ttnn (-p fake_ttnn_plugin)") + from tt_diffusion_planner.reference.weights import find_weights_dir, load_weights + + wd = find_weights_dir() + if wd is None: + pytest.skip("weights not found") + ttnn.reset() + undo = [_load("fake_ttnn_cnn").install(ttnn), _load("fake_ttnn_attention").install(ttnn)] + undo.append(_install_planner_ops(ttnn)) + from tt_diffusion_planner.ttaw.device import close_device, open_device + from tt_diffusion_planner.tt.model import TtDiffusionPlanner + + weights = load_weights(wd) + dev = open_device(0, dispatch="eth", num_command_queues=1) + tt = TtDiffusionPlanner(dev, weights, debug=True) + yield tt, weights + tt.release() + close_device(dev) + for fn in reversed(undo): + fn() + + +@pytest.fixture(scope="module") +def scene(fake_planner): + tt, weights = fake_planner + path = GOLDENS / f"{SCENE}.npz" + if not path.is_file(): + pytest.skip(f"research goldens {path} not found") + with np.load(path, allow_pickle=False) as z: + gold = {k: z[k] for k in z.files if k != "__meta__"} + raw = {k: gold[f"in.{k}"] for k in C.INPUT_NAMES} + return hp.prepare(raw, weights.normalization.observation), gold + + +def test_encoder_taps_match_the_reference(fake_planner, scene): + tt, _ = fake_planner + prep, gold = scene + taps = tt.encoder_taps(prep, eager=True) + for name, _ in C.TOKEN_LAYOUT: + rows = np.flatnonzero(gold[f"host.valid.{name}"]) + if rows.size: + assert pcc(taps[f"enc.{name}"][rows], gold[f"enc.{name}"][rows]) > 0.999, name + for c in ("ego", "neighbor", "lane", "route", "line_string"): + rows = gold[f"enc.{c}.pre.rows"] + assert pcc(taps[f"enc.{c}.pre"][rows], gold[f"enc.{c}.pre"]) > 0.9999, c + assert pcc(taps[f"enc.{c}.mixer"][rows], gold[f"enc.{c}.mixer"]) > 0.999, c + tok = np.flatnonzero(gold["host.token_valid"]) + assert pcc(taps["enc.encoding"][tok], gold["enc.encoding"][tok]) > 0.999 + + +def test_decode_once_matches_the_reference(fake_planner, scene): + tt, _ = fake_planner + prep, gold = scene + rows = gold["dec.rows"] + for k in (0, 10): + got = tt.decode_once(prep, gold["dec.x_in"][k], float(gold["dec.t"][k]), encoding=gold["enc.encoding"], + eager=True) + assert not got[:, 0].any() # the masked t = 0 columns + assert pcc(got[rows][:, 1:], gold["dec.out"][k][:, 1:]) > 0.999, k + + +def test_plan_and_trace(fake_planner, scene): + """The plan eagerly vs the reference's final x0 / logits, then captured: replay == eager.""" + tt, weights = fake_planner + prep, gold = scene + eager = tt.forward(prep, eager=True) + rows = gold["dec.rows"] + assert pcc(eager["final_x0"][rows], gold["final_x0"][rows]) > 0.999 + np.testing.assert_array_equal(eager["final_x0"][:, 0], prep.decoder.current_states) # prefix constraint + assert np.abs(eager["logit"] - gold["turn.logit"]).max() < 0.05 + assert len(eager["denoising_steps"]) == C.DPM_SOLVER_STEPS + 1 + np.testing.assert_array_equal(eager["denoising_steps"][-1][0], eager["final_x0"][0]) + tt.capture() + replay = tt.forward(prep) + for k in ("final_x0", "logit"): + np.testing.assert_array_equal(replay[k], eager[k]) + out = hp.make_output(replay["final_x0"], replay["logit"], prep, weights.normalization, + {k: v[3] for k, v in hp.RUNTIME_PARAMS.items()}) + assert int(out.turn_indicator["command"]) == int(gold["out.turn_command"]) diff --git a/code/tt_diffusion_planner/tests/test_tt_params_host.py b/code/tt_diffusion_planner/tests/test_tt_params_host.py new file mode 100644 index 0000000000000000000000000000000000000000..003a38a7cb9f2a2ac6a7edfc6d14681cd44cac0f --- /dev/null +++ b/code/tt_diffusion_planner/tests/test_tt_params_host.py @@ -0,0 +1,158 @@ +# SPDX-License-Identifier: Apache-2.0 +"""The exact rewrites of ``tt/params.py`` and the input packing of ``tt/inputs.py`` against the CPU reference +(float64, no device): every constant the ttnn graph holds is the reference's math rearranged. + + TT_VISIBLE_DEVICES=none python -m pytest -q -p fake_ttnn_plugin code/tt_diffusion_planner/tests/test_tt_params_host.py +""" +from __future__ import annotations + +import numpy as np +import pytest + +from tt_diffusion_planner.host import pipeline as hp +from tt_diffusion_planner.host.solver import apply_prefix_constraint, dpm_solver_sample +from tt_diffusion_planner.reference import config as C +from tt_diffusion_planner.tt import config as T +from tt_diffusion_planner.tt import inputs as I +from tt_diffusion_planner.tt import params as P + +torch = pytest.importorskip("torch") +F = torch.nn.functional +PKG = __import__("pathlib").Path(__file__).resolve().parents[1] + + +@pytest.fixture(scope="module") +def weights(): + from tt_diffusion_planner.reference.weights import find_weights_dir, load_weights + + wd = find_weights_dir() + if wd is None: + pytest.skip("Diffusion Planner v5.0 weights not found") + return load_weights(wd) + + +@pytest.fixture(scope="module") +def prepared(weights): + from tt_diffusion_planner.ttaw.io import load_named_arrays + + raw = load_named_arrays(str(PKG / "samples" / "kashiwanoha_dense.npz"), C.INPUT_SCHEMA) + return hp.prepare(raw, weights.normalization.observation) + + +def d(a): + return np.asarray(a, np.float64) + + +def test_lane_aux_is_the_speed_and_attribute_embeddings(weights): + p = weights.params + rng = np.random.default_rng(0) + for cat, mod in (("lane", "encoder.lane_encoder"), ("route", "encoder.route_encoder")): + e = 40 + speed = rng.random(e).astype(np.float32) * 2 + has = rng.random(e) < 0.6 + speed[~has] = 0.0 + attr = (rng.random((e, C.LANE_ATTRIBUTE_DIM)) < 0.2).astype(np.float32) + got = d(P.lane_aux_features(speed, has, attr)) @ d(P.lane_aux(p, cat)) + lin = d(speed)[:, None] @ d(p[f"{mod}.speed_limit_emb.w"]) + d(p[f"{mod}.speed_limit_emb.b"]) + want = np.where(has[:, None], lin, d(p[f"{mod}.unknown_speed_emb"])[None]) \ + + d(attr) @ d(p[f"{mod}.attribute_emb.w"]) + d(p[f"{mod}.attribute_emb.b"]) + np.testing.assert_allclose(got, want, rtol=1e-12, atol=1e-12) + + +def test_neighbor_aux_and_pos_aug(weights, prepared): + p = weights.params + f = prepared.features + t = d(f.neighbor_type) + got = np.concatenate([t, np.ones((t.shape[0], 1))], 1) @ d(P.neighbor_aux(p)) + want = t @ d(p["encoder.neighbor_encoder.type_emb.w"]) + d(p["encoder.neighbor_encoder.type_emb.b"]) + np.testing.assert_allclose(got, want, rtol=1e-12, atol=1e-12) + inp = I.plan_inputs(prepared) + pos = d(inp["pos_aug"][0, 0]) @ d(P.pos_aug(p)) + tv = np.asarray(f.token_valid, bool) + ref = (d(f.pos) @ d(p["encoder.pos_emb.w"]) + d(p["encoder.pos_emb.b"])) * tv[:, None] + np.testing.assert_allclose(pos[:T.TOKENS_REAL], ref, rtol=1e-12, atol=1e-12) + assert not pos[T.TOKENS_REAL:].any() + + +def test_agent_rows_and_masked_projection(weights): + p = weights.params + rows = P.agent_rows(p) + emb = d(p["decoder.dit.agent_embedding"]) + b2 = d(p["decoder.dit.preproj.fc2.b"]) + np.testing.assert_allclose(d(rows[0]), b2 + emb[0], rtol=1e-6) + np.testing.assert_allclose(d(rows[1:]), np.repeat((b2 + emb[1])[None], T.AGENTS - 1, 0), rtol=1e-6) + lin = P.final_projection_masked(p) + assert not lin.w[:, :4].any() and not lin.b[:4].any() + np.testing.assert_array_equal(lin.w[:, 4:], np.asarray(p["decoder.dit.final_layer.proj.4.w"])[:, 4:]) + + +def test_turn_weights_reproduce_the_head(weights): + from tt_diffusion_planner.reference.model import TurnHead, torch_params + + p = weights.params + rng = np.random.default_rng(1) + enc = rng.standard_normal((T.TOKENS_REAL, C.HIDDEN_DIM)) + x0 = rng.standard_normal((C.MAX_NUM_AGENTS, C.OUTPUT_T + 1, C.POSE_DIM)) + w = P.turn_weights(p) + got = x0[0].reshape(-1) @ d(w["w_sel"]) + enc.sum(0) @ d(w["w_pool"]) + d(w["b"]) + head = TurnHead(torch_params(p, torch.float64)) + want = head.forward(torch.from_numpy(enc), x0).numpy() + np.testing.assert_allclose(got, want, rtol=1e-6, atol=1e-6) + + +def test_step_tables_and_solver_recursion(weights): + """The y-recursion with (A, B, Cm) and the masked model output equals the node's DPM-Solver++(2M) with the + prefix constraint (dpm_solver_sample) on a deterministic stand-in model.""" + tab = P.step_tables(weights.params) + assert tab.nfe == C.DPM_SOLVER_STEPS + 1 and len(tab.solver) == C.DPM_SOLVER_STEPS + assert len(tab.blocks) == tab.nfe and len(tab.blocks[0]) == C.DIT_DEPTH + rng = np.random.default_rng(2) + agents = 5 + W = rng.standard_normal((324, 324)).astype(np.float32) * 0.05 + cs = rng.standard_normal((agents, 4)).astype(np.float32) + x_T = rng.standard_normal((agents, 81, 4)).astype(np.float32) + + def model(x, t): + return np.tanh(x.reshape(agents, -1) @ W + float(t)).reshape(agents, 81, 4).astype(np.float32) + + ref = dpm_solver_sample(x_T, model, lambda x: apply_prefix_constraint(x, cs)) + mask = np.ones((81, 4), np.float32) + mask[0] = 0 + y = (x_T * mask).astype(np.float64) + m_prev = None + iters = [] + for k in range(tab.nfe): + x = y + np.concatenate([cs[:, None], np.zeros((agents, 80, 4))], 1) + if k > 0: + iters.append(x) + m = model(x.astype(np.float32), tab.eval_times[k]).astype(np.float64) * mask + if k == tab.nfe - 1: + break + a, b, c = tab.solver[k] + y = a * y - b * m + (c * m_prev if m_prev is not None else 0.0) + m_prev = m + final = m + np.concatenate([cs[:, None], np.zeros((agents, 80, 4))], 1) + iters.append(final) + np.testing.assert_allclose(final, ref.final_x, rtol=1e-5, atol=1e-5) + for got, want in zip(iters, ref.denoising_steps): + np.testing.assert_allclose(got, want, rtol=1e-5, atol=1e-5) + assert np.allclose(tab.eval_times, ref.eval_times) + + +def test_plan_inputs_layout(prepared): + inp = I.plan_inputs(prepared) + assert set(inp) == set(I.INPUT_SPECS) + f = prepared.features + np.testing.assert_array_equal(inp["neighbor_x"][0], f.neighbor[:, 25:31]) + assert not f.neighbor[:, :25].any() and not f.ego[6:].any() # the island's zero rows + np.testing.assert_array_equal(inp["ego_x"][0, 0], f.ego[:6]) + row = inp["fusion_key_row"][0, 0, 0] + assert np.isneginf(row[T.TOKENS_REAL:]).all() and (row[:T.TOKENS_REAL][f.key_valid] == 0).all() + arow = inp["agent_key_row"][0, 0, 0] + assert np.isneginf(arow[C.MAX_NUM_AGENTS:]).all() and arow[0] == 0 + cs, y0 = inp["cs"][0, 0], inp["y0"][0, 0] + np.testing.assert_array_equal(cs[:C.MAX_NUM_AGENTS, :4], prepared.decoder.current_states) + assert not cs[:, 4:].any() and not y0[:, :4].any() and not cs[C.MAX_NUM_AGENTS:].any() + np.testing.assert_array_equal(y0[:C.MAX_NUM_AGENTS, 4:], prepared.x_T.reshape(C.MAX_NUM_AGENTS, -1)[:, 4:]) + warm = I.warmup_inputs() + assert all(warm[k].shape == v for k, v in I.INPUT_SPECS.items()) diff --git a/code/tt_diffusion_planner/tt/__init__.py b/code/tt_diffusion_planner/tt/__init__.py new file mode 100644 index 0000000000000000000000000000000000000000..7514d4b75dedc2b42f682ffd9c7d605a3a660298 --- /dev/null +++ b/code/tt_diffusion_planner/tt/__init__.py @@ -0,0 +1,27 @@ +# SPDX-License-Identifier: Apache-2.0 +"""tt-nn implementation of Diffusion Planner v5.0 for one Blackhole p150 (12x10 grid, ETH dispatch). + +Planned layout (PORT_LOG.md milestones M1-M5; nothing here yet at M0): + +- ``model.py`` ``TtDiffusionPlanner(device, weights, knobs=None, debug=False)``: the canonical weights of + ``reference.weights`` -> device tensors (bf16; the ego / neighbour pre-projection island in fp32, + pad-relative with the constants of ``reference.rewrites.island_constants``; the per-step adaLN tables + of ``reference.rewrites.adaln_tables`` folded into LayerNorm affine rows; the solver coefficients of + ``host.solver.solver_plan``), and the variants of a ``ttaw.trace.TraceRunner``: ``plan`` (encoder -> + fusion -> hoisted cross K|V -> 11 x (DiT + fp32 DPM-Solver++(2M) update + prefix constraint) -> turn + head, one packed readback), plus the debug variants ``encoder_taps`` and ``decode_once`` used by + ``tests/test_pcc_device.py``. ``forward(prepared) -> {"final_x0", "logit"[, "denoising_steps"]}``. +- ``encoder.py`` MLP-Mixer trunk (token mixing as transpose -> 2-D ``[1, 1, E*128, K] @ W`` -> transpose, probe P12), + entity heads, positional embedding, the fusion transformer (C20 SDPA 564 x 564 with the -inf key + mask). +- ``decoder.py`` one DiT evaluation with table-driven modulation, masked self-attention (321 x 321) and + cross-attention on the hoisted K|V (321 x 564); the solver update and prefix constraint. +- ``kernels/`` none planned for the functional port (PLAN.md 2.12: no custom kernels); fused mixer / DiT-step + kernels and the single-megakernel attempt are optimization work (PLAN.md 5, D19). + +Rules: read the grid from the device (``ttaw.device.compute_grid``); explicit ``compute_kernel_config`` on every +matmul (``ttaw.precision``; HiFi4 for fp32 operands); ``epsilon=1e-5`` on every LayerNorm; ``is_causal=False`` and +``scale=1/sqrt(32)`` on every SDPA (through the C20 wrapper ``ttaw.ops.attention``); no host reads inside a capture; +no program compiled after the first capture. ``api.py`` imports this package inside ``_build`` only, so +``import tt_diffusion_planner`` stays free of ttnn while the modules here may import it at the top. +""" diff --git a/code/tt_diffusion_planner/tt/config.py b/code/tt_diffusion_planner/tt/config.py new file mode 100644 index 0000000000000000000000000000000000000000..fa8e79259d92172535a8096bf6500b9c4e620aa4 --- /dev/null +++ b/code/tt_diffusion_planner/tt/config.py @@ -0,0 +1,111 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Device-side constants and the precision policy of the ttnn graph (COMPILE parameters, PLAN.md 0.4). + +Shapes. The masked attentions run on tile-aligned key counts: a user mask with an unaligned key count lets the +padded keys of the last tile into the softmax (``ttaw.ops.attention``, C20), so the fusion transformer runs on +:data:`TOKENS` = 576 tokens (564 + 12 zero pad tokens, masked as keys) and the decoder on :data:`AGENTS` = 352 agents +(321 + 31 extra rows, masked as self-attention keys). Both counts are what the tiles hold anyway (no extra compute). +Rows past the real counts are computed (finite) and never read back. The cross-attention keys are the real 564 +tokens (no mask: the op masks its own padding element by element). + +Solver state. ``y = x * mask0`` (the state with the t = 0 columns, the prefix-constrained current state, zeroed) is +the fp32 state of the on-device DPM-Solver++(2M) loop; the decoder input is ``x = y + cs`` (``cs`` = current +states in the t = 0 columns) and the last projection's t = 0 output columns are zeroed in the weights, so the +model output is ``m * mask0`` and every update ``y' = A y - B m_k + C m_(k-1)`` keeps ``y * mask0 = y``. This is +the node's update + ``apply_prefix_constraint`` exactly (``m``'s t = 0 slot is never used: the correction +overwrites it). + +Numerics options (:data:`KNOBS`, measured on the device, PORT_LOG.md 5.2): the fused ``ttnn.layer_norm`` loses +the per-entity signal on offset-dominated rows (fp32 decomposition for the mixers and the decoder: ``LN_FP32``); a +device fp32 matmul truncates its operands TF32-like, which the mixers amplify ~100x and which, compounded over the 11 +decoder evaluations, moves the plan (split hi / lo matmuls for the mixer inputs and every decoder linear: +``SPLIT_MATMUL``); and the bf16 SDPA kernel is the largest encoder error left on the sensitive nuScenes instants +(fp32 matmul attention in the fusion and the decoder: ``ATTN_MATMUL``). With the defaults the 27 e2e scenes (7 +research + 20 nuScenes instants) stay within ego max 0.31 m / mean 0.14 m / neighbour median-max 0.09 m of the fp32 +reference (gates 1.0 / 0.3 / 1.5 m); the defaults of the first device round failed the mean on one instant (0.35 m). + +Precision (:data:`DEFAULT_PRECISION`, ``ttaw.precision.PrecisionPolicy`` rules, first match wins; override with +``DIFFUSION_PLANNER_PRECISION="dec.*=HiFi2+fp32:a=bf16"``): every matmul / LayerNorm gets an explicit compute +config; ``w=`` is the weight dtype, ``a=`` the dtype of the module's residual stream (sub-layer outputs are produced +in it and LayerNorm keeps it). The ego / neighbour pre-projection island is fp32 and pad-relative (probe P12, a +device fp32 matmul is TF32-like); the decoder pre-projection reads the fp32 solver state with fp32 weights. +""" +from __future__ import annotations + +from typing import Dict + +from ..reference import config as C + +__all__ = ["TOKENS", "TOKENS_REAL", "AGENTS", "AGENTS_REAL", "STATE_COLS", "STATE_COLS_T0", "TILE", + "ISLAND_ROWS", "LANE_AUX_DIM", "NEIGHBOR_AUX_DIM", "POS_AUG_DIM", "DEFAULT_PRECISION", "MIXER_T", + "MIXER_CIN", "KNOBS", "globs"] + +TILE = 32 + + +def _aligned(n: int) -> int: + return -(-n // TILE) * TILE + + +TOKENS_REAL = C.ENCODING_TOKEN_NUM # 564 encoder tokens (the cross-attention keys) +TOKENS = _aligned(TOKENS_REAL) # 576: fusion rows / masked keys +AGENTS_REAL = C.MAX_NUM_AGENTS # 321 +AGENTS = _aligned(AGENTS_REAL) # 352: decoder rows / masked self-attention keys +STATE_COLS = C.DIT_INPUT_DIM # 324 = 81 points x 4 +STATE_COLS_T0 = C.POSE_DIM # columns 0..3 hold the t = 0 point (prefix constraint) + +# pre-projection island rows (the only non-zero time rows after the in-graph truncation, SPEC 3.8) +ISLAND_ROWS = {"ego": tuple(range(C.EGO_HISTORY_KEEP.start, C.EGO_HISTORY_KEEP.stop)), + "neighbor": tuple(range(C.NEIGHBOR_HISTORY_KEEP.start, C.NEIGHBOR_HISTORY_KEEP.stop))} +# mixer token axis length (T) and input channels per category (host features) +MIXER_T = {"ego": len(ISLAND_ROWS["ego"]), "neighbor": len(ISLAND_ROWS["neighbor"]), "lane": C.POINTS_PER_SEGMENT, + "route": C.POINTS_PER_SEGMENT, "polygon": C.POINTS_PER_POLYGON, "line_string": C.POINTS_PER_LINE_STRING} +MIXER_CIN = {"ego": C.POSE_DIM, "neighbor": C.NEIGHBOR_FEATURE_DIM, "lane": C.LANE_FEATURE_DIM, + "route": C.LANE_FEATURE_DIM, "polygon": C.POLYGON_FEATURE_DIM, + "line_string": C.LINE_STRING_FEATURE_DIM} +# exact affine rewrites of the small embeddings (tt.params): the host builds the input columns +NEIGHBOR_AUX_DIM = 3 + 1 # [type one-hot (3), 1] @ [W_type; b_type] +LANE_AUX_DIM = 3 + C.LANE_ATTRIBUTE_DIM + 1 # [speed*has, has, 1-has, attributes (25), 1] @ [w_s; b_s; unk; W_a; b_a] +POS_AUG_DIM = C.POS_FEATURE_DIM + 1 # [pos (14), token valid] @ [W_pos; b_pos] + +DEFAULT_PRECISION: Dict[str, str] = { + "enc.island.*": "HiFi4+fp32:w=fp32:a=fp32", + "enc.mixer.*": "HiFi4+fp32:w=bf16:a=fp32", + "enc.fusion*": "HiFi4+fp32:w=bf16:a=fp32", + "enc.*": "HiFi4+fp32:w=bf16:a=fp32", + "dec.preproj.fc1": "HiFi4+fp32:w=fp32:a=fp32", + "dec.*": "HiFi4+fp32:w=bf16:a=fp32", + "turn": "HiFi4+fp32:w=fp32:a=fp32", +} + + +def _knobs(): + from ..ttaw.knobs import Knob, Knobs + + return Knobs("DIFFUSION_PLANNER", [ + Knob("LN_FP32", "enc.mixer.*,dec.*", "comma-separated module globs whose LayerNorms run as an fp32 " + "decomposition (mean, variance, rsqrt as fp32 SFPU ops) instead of ttnn.layer_norm, whose error on " + "offset-dominated rows (|mean| / std 10-30) is rel-L2 0.03 (PORT_LOG 5.2); 'none' = empty"), + Knob("HIDDEN_FP32", "", "comma-separated module globs whose hidden MLP activations are fp32 (default " + "bf16); 'none' = empty"), + Knob("SPLIT_MATMUL", "enc.island.*,enc.pre.*,dec.*", "comma-separated module globs whose fp32 " + "matmuls are split into bf16 hi / fp32 lo parts (2-3 matmuls, ~1e-5 relative instead of the TF32-like " + "~1e-3; their hidden activations stay fp32): the mixer inputs and every decoder linear; 'none' = " + "empty"), + Knob("ATTN_FP32_ACC", "", "comma-separated module globs whose SDPA runs with fp32 accumulation (probe P7: " + "half the max error, ~1.4x the time); 'none' = empty"), + Knob("ATTN_MATMUL", "enc.fusion.attn,dec.*", "comma-separated module globs (enc.fusion.attn, " + "dec.self_attn, dec.cross_attn) whose attention runs as fp32 matmuls + softmax (C20 attention_matmul) " + "instead of the bf16 SDPA kernel; 'none' = empty"), + ]) + + +KNOBS = _knobs() + + +def globs(text: str): + """``"a.*, b"`` -> ``("a.*", "b")``; ``""`` / ``"none"`` -> ``()``.""" + t = (text or "").strip() + if t.lower() in ("", "none"): + return () + return tuple(g.strip() for g in t.split(",") if g.strip()) diff --git a/code/tt_diffusion_planner/tt/decoder.py b/code/tt_diffusion_planner/tt/decoder.py new file mode 100644 index 0000000000000000000000000000000000000000..36660eed1e7ac6463b5df88cb6d02b83e6ddcef5 --- /dev/null +++ b/code/tt_diffusion_planner/tt/decoder.py @@ -0,0 +1,183 @@ +# SPDX-License-Identifier: Apache-2.0 +"""The DiT decoder, the on-device DPM-Solver++(2M) loop and the turn head (SPEC 4.3-4.4, 5.1; ``dpm_solver.cpp``). + +One evaluation (``T4M dit.py``; rows = the 352 decoder agents of ``tt/config.py``): + +- ``preproj``: ``x [352, 324] (fp32 state) -> 512 (exact GELU) -> 256`` + agent embedding rows (folded bias rows); +- 3 DiT blocks with the per-step adaLN tables **folded into the LayerNorm affine** (``reference.rewrites``): + ``x += gate_msa * out(attn_masked(qkv(LN1_k(x))))``, ``x += gate_mlp * mlp1(LN2_k(x))`` (tanh GELU), + ``x += out(attn(q(LN3(x)), K_i, V_i))`` on the hoisted cross K / V, ``x += mlp2(LN4(x))``; +- final layer: ``LN_k`` (folded) -> ``LN(256)`` -> ``1024`` (tanh GELU) -> ``LN(1024)`` -> ``324`` with the t = 0 output + columns zeroed (``m * mask0``, fp32). + +Solver (``tt/config.py``): state ``y = x * mask0``; evaluation ``k`` reads ``x = y + cs``; updates +``y' = A_k y - B_k m_k + Cm_k m_(k-1)`` (``tt.params.solver_coefficients``); after the 11th evaluation (t = 1e-3, +denoise to zero) ``final_x0 = m_10 + cs``. All state and model outputs are fp32; the scalars are constants of the +unrolled loop (one trace per plan). + +Attention (``ATTN_MATMUL`` knob, default ``dec.*``): fp32 Q / K / V, two fp32 matmuls + softmax (C20 +``attention_matmul``); the cross-attention then reads the 576 encoder rows with the 12 pad tokens masked. With the +knob off: the bf16 C20 ``sdpa`` (self-attention masked on 352 keys; cross-attention on the 564 real tokens, no mask). +Every linear is a split matmul by default (``SPLIT_MATMUL`` ``dec.*``: fp32-accurate, PORT_LOG decision 11). + +Turn head: ``logit = final_x0[0] @ W_sel + (ones_564 @ encoding) @ (W_pool / 564) + b`` (fp32). +""" +from __future__ import annotations + +from typing import Any, Dict, List, Mapping, Sequence, Tuple + +import numpy as np + +from ..reference import config as C +from ..ttaw.ops import attention as A +from . import config as T +from . import params as P +from .layers import ATTN, Build, Const, LayerNorm, make_linear + +__all__ = ["TtDecoder", "TtTurnHead", "STEP_KEYS"] + +STEP_KEYS = ("n1_g", "n1_b", "gate_msa", "n2_g", "n2_b", "gate_mlp") + + +class TtDecoder: + """DiT weights, the per-step tables as device rows, and the evaluation / solver graph.""" + + def __init__(self, build: Build, p: Mapping[str, np.ndarray], tables: P.StepTables): + self.build = build + self.tables = tables + D = "dec.block" + self.stream = build.stream(D) + # the pre-projection reads the fp32 solver state: with split matmuls (SPLIT_MATMUL "dec.preproj.*") the + # positions are not truncated to the TF32-like operand of a device fp32 matmul (probe P12) + self.pre1 = make_linear(build, P.linear(p, "decoder.dit.preproj.fc1"), "dec.preproj.fc1", + out=build.hidden("dec.preproj.fc1"), activation="gelu") + self.pre2 = make_linear(build, P.linear(p, "decoder.dit.preproj.fc2", bias=False), "dec.preproj.fc2", + out=self.stream) + self.agent_rows = Const(build, P.agent_rows(p, T.AGENTS), self.stream) + self.blocks = [] + for i in range(C.DIT_DEPTH): + B = f"decoder.dit.blocks.{i}" + self.blocks.append({ + "ln": LayerNorm(build, None, D), # affine rows per step (folded adaLN) + "qkv": make_linear(build, P.linear(p, f"{B}.attn.qkv"), D, + out="float32" if build.attn_matmul("dec.self_attn") else ATTN), + "out": make_linear(build, P.linear(p, f"{B}.attn.out"), D, out=self.stream), + "m1a": make_linear(build, P.linear(p, f"{B}.mlp1.fc1"), D, out=build.hidden(D), + activation="gelu_tanh"), + "m1b": make_linear(build, P.linear(p, f"{B}.mlp1.fc2"), D, out=self.stream), + "n3": LayerNorm(build, P.norm(p, f"{B}.norm3"), D), + "cq": make_linear(build, P.linear(p, f"{B}.cross_attn.q"), D, + out="float32" if build.attn_matmul("dec.cross_attn") else ATTN), + "cout": make_linear(build, P.linear(p, f"{B}.cross_attn.out"), D, out=self.stream), + "n4": LayerNorm(build, P.norm(p, f"{B}.norm4"), D), + "m2a": make_linear(build, P.linear(p, f"{B}.mlp2.fc1"), D, out=build.hidden(D), + activation="gelu_tanh"), + "m2b": make_linear(build, P.linear(p, f"{B}.mlp2.fc2"), D, out=self.stream)}) + kvw = [P.linear(p, f"decoder.dit.blocks.{i}.cross_attn.kv") for i in range(C.DIT_DEPTH)] + H = C.HIDDEN_DIM + kv_out = "float32" if build.attn_matmul("dec.cross_attn") else ATTN + self.k_lin = [make_linear(build, P.Lin(lin.w[:, :H], lin.b[:H]), "dec.kv", out=kv_out) for lin in kvw] + self.v_lin = [make_linear(build, P.Lin(lin.w[:, H:], lin.b[H:]), "dec.kv", out=kv_out) for lin in kvw] + F = "dec.final" + self.fin_ln = LayerNorm(build, None, F) + self.p0 = LayerNorm(build, P.norm(p, "decoder.dit.final_layer.proj.0"), F) + self.p1 = make_linear(build, P.linear(p, "decoder.dit.final_layer.proj.1"), F, out=build.hidden(F), + activation="gelu_tanh") + self.p3 = LayerNorm(build, P.norm(p, "decoder.dit.final_layer.proj.3"), F) + self.p4 = make_linear(build, P.final_projection_masked(p), F, out="float32") + # per-step rows (RT-dev constants of the unrolled loop): [k][i][key] / [k]["g" | "b"] + self.step_rows = [self.upload_step(tables, k) for k in range(tables.nfe)] + self.scale = float(C.ATTN_SCALE) + self.attn_fp32 = {"self": build.attn_fp32_acc("dec.self_attn"), "cross": build.attn_fp32_acc("dec.cross_attn")} + self.attn_mm = {"self": build.attn_matmul("dec.self_attn"), "cross": build.attn_matmul("dec.cross_attn")} + if self.attn_mm["cross"]: # fp32 cross-attention over the 576 encoder rows, the 12 pad tokens masked + row = np.where(np.arange(T.TOKENS) < T.TOKENS_REAL, 0.0, -np.inf).astype(np.float32) + self.cross_mask = Const(build, np.broadcast_to(row, (T.AGENTS, T.TOKENS)), "bfloat16") + + def upload_step(self, tables: P.StepTables, k: int) -> Dict[str, Any]: + b = self.build + return {"blocks": [{key: b.upload(np.asarray(blk[key]).reshape(1, -1), "float32") for key in STEP_KEYS} + for blk in tables.blocks[k]], + "g": b.upload(np.asarray(tables.final[k]["g"]).reshape(1, -1), "float32"), + "b": b.upload(np.asarray(tables.final[k]["b"]).reshape(1, -1), "float32")} + + # ------------------------------------------------------------------------------------------------------------ + def cross_kv(self, enc) -> List[Tuple[Any, Any]]: + """Hoisted cross-attention K / V heads of the three blocks (once per plan) from the encoding ``[1, 1, 576, + 256]``: ``[1, 8, 564, 32]`` (the 564 real tokens; SDPA masks its own tile padding) or, with the fp32 matmul + attention, ``[1, 8, 576, 32]`` fp32 (the pad tokens are masked by ``cross_mask``).""" + import ttnn + + if not self.attn_mm["cross"] and int(enc.shape[-2]) != T.TOKENS_REAL: + enc = ttnn.slice(enc, [0, 0, 0, 0], [1, 1, T.TOKENS_REAL, C.HIDDEN_DIM]) + return [(A.split_heads(k(enc), C.NUM_HEADS), A.split_heads(v(enc), C.NUM_HEADS)) + for k, v in zip(self.k_lin, self.v_lin)] + + def _attend(self, kind: str, q, k, v, mask): + """Attention of ``kind`` ("self" / "cross") -> concatenated heads ``[1, 1, 352, 256]``.""" + if self.attn_mm[kind]: + return A.merge_heads(A.attention_matmul(q, k, v, scale=self.scale, attn_mask=mask)) + return A.sdpa(q, k, v, scale=self.scale, attn_mask=mask, concat_heads=True, fp32_acc=self.attn_fp32[kind]) + + def evaluate(self, x, rows: Mapping[str, Any], kv: Sequence[Tuple[Any, Any]], self_mask): + """One decoder evaluation: ``x`` ``[1, 1, 352, 324]`` fp32 (prefix-constrained) -> ``m * mask0`` fp32.""" + import ttnn + + h = ttnn.add(self.pre2(self.pre1(x)), self.agent_rows()) # [1, 1, 352, 256] + for b, r, (kc, vc) in zip(self.blocks, rows["blocks"], kv): + q, k, v = A.split_qkv(b["qkv"](b["ln"](h, r["n1_g"], r["n1_b"])), C.NUM_HEADS) + a = b["out"](self._attend("self", q, k, v, self_mask)) + h = ttnn.add(h, ttnn.multiply(a, r["gate_msa"])) + m = b["m1b"](b["m1a"](b["ln"](h, r["n2_g"], r["n2_b"]))) + h = ttnn.add(h, ttnn.multiply(m, r["gate_mlp"])) + qc = A.split_heads(b["cq"](b["n3"](h)), C.NUM_HEADS) + cmask = self.cross_mask() if self.attn_mm["cross"] else None + h = ttnn.add(h, b["cout"](self._attend("cross", qc, kc, vc, cmask))) + h = ttnn.add(h, b["m2b"](b["m2a"](b["n4"](h)))) + f = self.fin_ln(h, rows["g"], rows["b"]) + return self.p4(self.p3(self.p1(self.p0(f)))) + + def solve(self, y0, cs, kv, self_mask, *, ego_steps: bool = True): + """The 11 evaluations with the DPM-Solver++(2M) updates. Returns ``(final_x0, ego_rows)``: ``final_x0`` + ``[1, 1, 352, 324]`` fp32 and, with ``ego_steps``, the ego row of the 11 published iterates + ``[1, 1, 11, 324]`` (``~/debug/denoising_steps``; ``None`` otherwise).""" + import ttnn + + y, m_prev, ego = y0, None, [] + for k in range(self.tables.nfe): + x = ttnn.add(y, cs) # prefix constraint + if ego_steps and k > 0: + ego.append(ttnn.slice(x, [0, 0, 0, 0], [1, 1, 1, T.STATE_COLS])) + m = self.evaluate(x, self.step_rows[k], kv, self_mask) + if k == self.tables.nfe - 1: + break + a_k, b_k, c_k = self.tables.solver[k] + y = ttnn.subtract(ttnn.multiply(y, a_k), ttnn.multiply(m, b_k)) + if m_prev is not None: + y = ttnn.add(y, ttnn.multiply(m_prev, c_k)) + m_prev = m + final = ttnn.add(m, cs) + if not ego_steps: + return final, None + ego.append(ttnn.slice(final, [0, 0, 0, 0], [1, 1, 1, T.STATE_COLS])) + return final, ttnn.concat(ego, dim=2) + + +class TtTurnHead: + """``logit [1, 1, 1, 5]`` from ``final_x0`` (row 0) and the token sum of the 564 real encoding rows.""" + + def __init__(self, build: Build, p: Mapping[str, np.ndarray]): + w = P.turn_weights(p) + self.sel = make_linear(build, P.Lin(w["w_sel"], w["b"]), "turn", out="float32") + self.pool = make_linear(build, P.Lin(w["w_pool"], None), "turn", out="float32") + ones = np.zeros((1, T.TOKENS), np.float32) + ones[0, :T.TOKENS_REAL] = 1.0 + self.ones = Const(build, ones, "bfloat16") + self.cfg = build.cfg("turn") + + def __call__(self, final_x0, encoding): + import ttnn + + row0 = ttnn.slice(final_x0, [0, 0, 0, 0], [1, 1, 1, T.STATE_COLS]) + s = ttnn.matmul(self.ones(), encoding, dtype=ttnn.float32, compute_kernel_config=self.cfg) # [1,1,1,256] + return ttnn.add(self.sel(row0), self.pool(s)) diff --git a/code/tt_diffusion_planner/tt/encoder.py b/code/tt_diffusion_planner/tt/encoder.py new file mode 100644 index 0000000000000000000000000000000000000000..60113640f3f4429dc0b5a76454809a9adf83b8e2 --- /dev/null +++ b/code/tt_diffusion_planner/tt/encoder.py @@ -0,0 +1,263 @@ +# SPDX-License-Identifier: Apache-2.0 +"""The encoder graph on the device: six MLP-Mixer trunks, the entity heads, the token assembly and the fusion +transformer -> ``encoding`` ``[1, 1, 576, 256]`` (564 real tokens + 12 zero pad tokens, ``tt/config.py``). + +Inputs are the persistent trace inputs written by :mod:`.inputs` (host features of ``host.features``, fp32 TILE): +``ego_x`` ``[1, 1, 6, 4]`` and ``neighbor_x`` ``[1, 320, 6, 9]`` (the 6 time rows the export keeps), the other +mixer inputs ``[1, E, T, C]``, the aux columns of the exact embedding rewrites (``tt/params.py``), the small-encoder +rows, ``token_valid`` ``[1, 1, 576, 1]``, ``pos_aug`` ``[1, 1, 576, 15]`` and the fusion key-bias row. + +Mixer trunk (``T4M mixer.py``; SPEC 4.2): ``channel_pre`` (C -> 128 -> 128) and ``token_pre`` over the T axis +(T -> 64 -> 64), then 6 blocks ``x += tokens_mlp(LN1(x)^T)^T; x += channels_mlp(LN2(x))`` on ``[1, E, 64, 128]``, +token mean -> ``[1, 1, E, 128]``. Token mixing is ``transpose -> 2-D [1, 1, E*128, 64] @ W -> transpose`` (probe +P12: 108 us at E = 320, never the broadcast-left batched matmul). Ego and neighbour run ``channel_pre`` + +``token_pre`` as the fp32 **pad-relative island** (``reference.rewrites``, P12): only deviations from the all-zero +agent pass the TF32-like matmuls. + +Fusion block (``encoder.py:327-333``): ``K|V = W_kv x`` from the un-normalised stream, ``Q = W_q LN1(x)``, masked +attention over 576 keys (the -inf key row expanded once per plan), ``x += out(.)``; ``x += mlp(LN2(x))``; final LN. +The attention is fp32 matmuls + softmax by default (C20 ``attention_matmul``, ``ATTN_MATMUL`` ``enc.fusion.attn``: the +bf16 SDPA was the largest encoder error on the sensitive nuScenes instants, PORT_LOG decision 11), else C20 ``sdpa``. +""" +from __future__ import annotations + +from typing import Any, Dict, Mapping, Optional + +import numpy as np + +from ..reference import config as C +from ..ttaw.ops import attention as A +from . import config as T +from . import params as P +from .layers import ATTN, Build, Const, LayerNorm, make_linear + +__all__ = ["MixerTrunk", "TtEncoder", "MIXER_CATS"] + +MIXER_CATS = ("ego", "neighbor", "lane", "route", "polygon", "line_string") +ENTITIES = dict(C.TOKEN_LAYOUT) + + +def _gelu(x): + import ttnn + + return ttnn.gelu(x, fast_and_approximate_mode=False) + + +class MixerTrunk: + """One MLP-Mixer category up to the token mean (``[1, 1, E, 128]``).""" + + def __init__(self, build: Build, p: Mapping[str, np.ndarray], cat: str): + self.cat, self.E, self.T = cat, ENTITIES[cat], T.MIXER_T[cat] + self.island = cat in T.ISLAND_ROWS + mods = P.mixer_module(p, cat) + mix = f"enc.mixer.{cat}" + self.stream = build.stream(mix) + if self.island: + isl = f"enc.island.{cat}" + w = P.island(p, cat) + self.isl_dtype = build.stream(isl) + + def lin(weights, act=None): + return make_linear(build, weights, isl, out=self.isl_dtype, activation=act) + + self.c1 = lin(P.Lin(w["c1_w"], w["c1_b"]), "gelu") + self.gelu_b1 = Const(build, w["gelu_b1"], self.isl_dtype) + self.c2 = lin(P.Lin(w["c2_w"], None)) + self.t1 = lin(P.Lin(w["t1_w"], None)) + self.t1_pad = Const(build, w["t1_pad"], self.isl_dtype) + self.g_pad = Const(build, w["g_pad"], self.isl_dtype) + self.t2 = lin(P.Lin(w["t2_w"], None)) + self.t2_pad = Const(build, w["t2_pad"], self.isl_dtype) + else: + pre = f"enc.pre.{cat}" + self.c1 = make_linear(build, mods["c1"], pre, out=build.hidden(pre), activation="gelu") + self.c2 = make_linear(build, mods["c2"], pre, out=build.hidden(pre)) + self.t1 = make_linear(build, mods["t1"], pre, out=build.hidden(pre), activation="gelu") + self.t2 = make_linear(build, mods["t2"], pre, out=self.stream) + self.blocks = [] + for blk in mods["blocks"]: + self.blocks.append({ + "n1": LayerNorm(build, blk["n1"], mix), + "tk1": make_linear(build, blk["tk1"], mix, out=build.hidden(mix), activation="gelu"), + "tk2": make_linear(build, blk["tk2"], mix, out=self.stream), + "n2": LayerNorm(build, blk["n2"], mix), + "ch1": make_linear(build, blk["ch1"], mix, out=build.hidden(mix), activation="gelu"), + "ch2": make_linear(build, blk["ch2"], mix, out=self.stream)}) + + # ------------------------------------------------------------------------------------------------------------ + def _tokens_view(self, x): + """``[1, E, 128, 64]`` <-> ``[1, 1, E*128, 64]`` (free: 128 rows are whole tiles).""" + import ttnn + + return ttnn.reshape(x, (1, 1, self.E * C.MIXER_CHANNELS, x.shape[-1])) + + def _entities_view(self, x): + import ttnn + + return ttnn.reshape(x, (1, self.E, C.MIXER_CHANNELS, x.shape[-1])) + + def pre(self, x): + """``[1, E, T, C]`` -> ``x0`` ``[1, E, 64, 128]`` (the ``enc..pre`` tap).""" + import ttnn + + if self.island: + h = ttnn.subtract(self.c1(x), self.gelu_b1()) # gelu(x W1 + b1) - gelu(b1) + dz = self.c2(h) # z - c0 [1, E, 6, 128] + dzt = self._tokens_view(ttnn.transpose(dz, -2, -1)) # [1, 1, E*128, 6] + t1 = ttnn.add(self._entities_view(self.t1(dzt)), self.t1_pad()) # t1_pad + dz^T W_t1[rows] + g = ttnn.subtract(_gelu(t1), self.g_pad()) + t2 = ttnn.add(self._entities_view(self.t2(self._tokens_view(g))), self.t2_pad()) + x0 = ttnn.transpose(t2, -2, -1) # [1, E, 64, 128] + if self.isl_dtype != self.stream: + x0 = ttnn.typecast(x0, self.stream_ttnn) + return x0 + z = self.c2(self.c1(x)) # [1, E, T, 128] + zt = self._tokens_view(ttnn.transpose(z, -2, -1)) # [1, 1, E*128, T] + t = self._entities_view(self.t2(self.t1(zt))) # [1, E, 128, 64] + x0 = ttnn.transpose(t, -2, -1) + if x0.dtype != self.stream_ttnn: + x0 = ttnn.typecast(x0, self.stream_ttnn) + return x0 + + @property + def stream_ttnn(self): + from ..ttaw.tensors import ttnn_dtype + + return ttnn_dtype(self.stream) + + def mix(self, x): + """The 6 MixerBlocks on ``[1, E, 64, 128]`` (the ``enc..mixer`` tap).""" + import ttnn + + for b in self.blocks: + y = self._tokens_view(ttnn.transpose(b["n1"](x), -2, -1)) # LN over the 128 channels, then T + y = ttnn.transpose(self._entities_view(b["tk2"](b["tk1"](y))), -2, -1) + x = ttnn.add(x, y) + x = ttnn.add(x, b["ch2"](b["ch1"](b["n2"](x)))) + return x + + def pool(self, x): + """Mean over the 64 tokens -> ``[1, 1, E, 128]``.""" + import ttnn + + m = ttnn.mean(x, dim=2, keepdim=True) # [1, E, 1, 128] + return ttnn.reshape(m, (1, 1, self.E, C.MIXER_CHANNELS)) + + +class _Head: + """``(m [+ aux @ W_aux]) -> LN(128) -> emb_project (128 -> 256 -> 256)`` (+ the route position embedding).""" + + def __init__(self, build: Build, p: Mapping[str, np.ndarray], cat: str, out_dtype: str, *, mixer=True): + mod = f"enc.head.{cat}" + stream = build.stream(mod) + m = P.mixer_module(p, cat) if mixer else P.small_module(p, cat) + self.aux = None + if cat == "neighbor": + self.aux = make_linear(build, P.Lin(P.neighbor_aux(p), None), mod, out=stream) + elif cat in ("lane", "route"): + self.aux = make_linear(build, P.Lin(P.lane_aux(p, cat), None), mod, out=stream) + self.norm = LayerNorm(build, m["norm"], mod) + self.e1 = make_linear(build, m["e1"], mod, out=build.hidden(mod), activation="gelu") + self.e2 = make_linear(build, m["e2"], mod, out=out_dtype) + self.route_pos = Const(build, p["encoder.route_position_embedding"], out_dtype) if cat == "route" else None + + def __call__(self, m, aux=None): + import ttnn + + if self.aux is not None: + m = ttnn.add(m, self.aux(aux)) + out = self.e2(self.e1(self.norm(m))) + if self.route_pos is not None: + out = ttnn.add(out, self.route_pos()) + return out + + +class _Small: + """goal / ego-shape / turn encoders: channel MLP (C -> 128 -> 128) -> LN -> emb_project.""" + + def __init__(self, build: Build, p: Mapping[str, np.ndarray], cat: str, out_dtype: str): + mod = f"enc.head.{cat}" + m = P.small_module(p, cat) + self.c1 = make_linear(build, m["c1"], mod, out=build.hidden(mod), activation="gelu") + self.c2 = make_linear(build, m["c2"], mod, out=build.stream(mod)) + self.head = _Head(build, p, cat, out_dtype, mixer=False) + + def __call__(self, x): + return self.head(self.c2(self.c1(x))) + + +class TtEncoder: + """Encoder + fusion on the device. ``forward(ctx, taps)`` -> ``encoding`` ``[1, 1, 576, 256]``; ``taps`` (a dict) + collects the device tensors of the ``reference.model.TAP_NAMES`` encoder taps when given.""" + + INPUTS = ("ego_x", "neighbor_x", "neighbor_aux", "static_x", "lane_x", "lane_aux", "route_x", "route_aux", + "polygon_x", "line_string_x", "goal_x", "ego_shape_x", "turn_x", "token_valid", "pos_aug", + "fusion_key_row") + + def __init__(self, build: Build, p: Mapping[str, np.ndarray]): + self.build = build + self.fstream = build.stream("enc.fusion") + self.trunks = {cat: MixerTrunk(build, p, cat) for cat in MIXER_CATS} + self.heads = {cat: _Head(build, p, cat, self.fstream) for cat in MIXER_CATS} + st = "enc.head.static" + self.static1 = make_linear(build, P.linear(p, "encoder.static_encoder.projection.fc1"), st, + out=build.hidden(st), activation="gelu") + self.static2 = make_linear(build, P.linear(p, "encoder.static_encoder.projection.fc2"), st, + out=self.fstream) + self.small = {cat: _Small(build, p, cat, self.fstream) for cat in ("goal", "ego_shape", "turn")} + self.pos = make_linear(build, P.Lin(P.pos_aug(p), None), "enc.tokens", out=self.fstream) + self.pad_tokens = Const(build, np.zeros((T.TOKENS - T.TOKENS_REAL, C.HIDDEN_DIM), np.float32), self.fstream) + fu = "enc.fusion" + self.attn_mm = build.attn_matmul("enc.fusion.attn") # fp32 matmul attention instead of bf16 SDPA + qkv_out = "float32" if self.attn_mm else ATTN + self.blocks = [] + for i in range(C.FUSION_DEPTH): + B = f"encoder.fusion.blocks.{i}" + self.blocks.append({ + "kv": make_linear(build, P.linear(p, f"{B}.attn.kv"), fu, out=qkv_out), + "n1": LayerNorm(build, P.norm(p, f"{B}.norm1"), fu), + "q": make_linear(build, P.linear(p, f"{B}.attn.q"), fu, out=qkv_out), + "out": make_linear(build, P.linear(p, f"{B}.attn.out"), fu, out=self.fstream), + "n2": LayerNorm(build, P.norm(p, f"{B}.norm2"), fu), + "fc1": make_linear(build, P.linear(p, f"{B}.mlp.fc1"), fu, out=build.hidden(fu), + activation="gelu"), + "fc2": make_linear(build, P.linear(p, f"{B}.mlp.fc2"), fu, out=self.fstream)}) + self.final_norm = LayerNorm(build, P.norm(p, "encoder.fusion.norm"), fu) + self.scale = float(C.ATTN_SCALE) + self.attn_fp32 = build.attn_fp32_acc("enc.fusion.attn") + + def forward(self, ctx: Mapping[str, Any], taps: Optional[Dict[str, Any]] = None): + import ttnn + + def tap(name, t): + if taps is not None: + taps[name] = t + return t + + out: Dict[str, Any] = {} + for cat in MIXER_CATS: + tr = self.trunks[cat] + x0 = tap(f"enc.{cat}.pre", tr.pre(ctx[f"{cat}_x"])) + x = tap(f"enc.{cat}.mixer", tr.mix(x0)) + aux = ctx.get(f"{cat}_aux") if cat in ("neighbor", "lane", "route") else None + out[cat] = self.heads[cat](tr.pool(x), aux) + out["static"] = self.static2(self.static1(ctx["static_x"])) + for cat, enc in self.small.items(): + out[cat] = enc(ctx[f"{cat}_x"]) + for name, _ in C.TOKEN_LAYOUT: + tap(f"enc.{name}", out[name]) + x = ttnn.concat([out[name] for name, _ in C.TOKEN_LAYOUT] + [self.pad_tokens()], dim=2) # [1,1,576,256] + x = ttnn.multiply(x, ctx["token_valid"]) # invalid entities -> 0 + x = tap("enc.tokens", ttnn.add(x, self.pos(ctx["pos_aug"]))) # + valid * (pos W + b) + mask = A.expand_key_bias(ctx["fusion_key_row"], T.TOKENS) # [1, 1, 576, 576], once per plan + for i, b in enumerate(self.blocks): + kv = b["kv"](x) # K | V from the un-normalised x + q, k, v = A.split_q_kv(b["q"](b["n1"](x)), kv, C.NUM_HEADS) + if self.attn_mm: + a = A.merge_heads(A.attention_matmul(q, k, v, scale=self.scale, attn_mask=mask)) + else: + a = A.sdpa(q, k, v, scale=self.scale, attn_mask=mask, concat_heads=True, fp32_acc=self.attn_fp32) + x = ttnn.add(x, b["out"](a)) + x = ttnn.add(x, b["fc2"](b["fc1"](b["n2"](x)))) + tap(f"enc.fusion.{i}", x) + return tap("enc.encoding", self.final_norm(x)) diff --git a/code/tt_diffusion_planner/tt/inputs.py b/code/tt_diffusion_planner/tt/inputs.py new file mode 100644 index 0000000000000000000000000000000000000000..3b51b9fa0c2bb7a0d248cf919ec5e1dec1066e35 --- /dev/null +++ b/code/tt_diffusion_planner/tt/inputs.py @@ -0,0 +1,102 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Host packing of one plan's :class:`host.pipeline.Prepared` into the persistent trace inputs (numpy, no ttnn). + +Every array is float32 in the 4-D shape of its device input (TILE, uploaded each plan): the mixer inputs keep only +the time rows the export reads (ego 0..5, neighbours 25..30: the pad-relative island needs nothing else), the aux +columns of the exact embedding rewrites (``tt/params.py``), the 576-row token arrays (564 real + 12 zero pad +tokens), the key-bias rows (0 / -inf; extra keys masked) and the solver's ``cs`` / ``y0`` on 352 rows. +""" +from __future__ import annotations + +from typing import Dict, Optional + +import numpy as np + +from ..host.pipeline import Prepared +from ..reference import config as C +from ..ttaw.ops.attention import key_bias_row +from . import config as T +from .params import lane_aux_features, pad_rows + +__all__ = ["plan_inputs", "INPUT_SPECS", "warmup_inputs", "decoder_state", "BF16_INPUTS"] + +f32 = np.float32 + +# name -> device shape (all fp32 TILE except the bf16 key-bias rows) +INPUT_SPECS: Dict[str, tuple] = { + "ego_x": (1, 1, T.MIXER_T["ego"], T.MIXER_CIN["ego"]), + "neighbor_x": (1, C.MAX_NUM_NEIGHBORS, T.MIXER_T["neighbor"], T.MIXER_CIN["neighbor"]), + "neighbor_aux": (1, 1, C.MAX_NUM_NEIGHBORS, T.NEIGHBOR_AUX_DIM), + "static_x": (1, 1, C.NUM_STATIC_OBJECTS, C.STATIC_OBJECT_DIM), + "lane_x": (1, C.NUM_SEGMENTS_IN_LANE, T.MIXER_T["lane"], T.MIXER_CIN["lane"]), + "lane_aux": (1, 1, C.NUM_SEGMENTS_IN_LANE, T.LANE_AUX_DIM), + "route_x": (1, C.NUM_SEGMENTS_IN_ROUTE, T.MIXER_T["route"], T.MIXER_CIN["route"]), + "route_aux": (1, 1, C.NUM_SEGMENTS_IN_ROUTE, T.LANE_AUX_DIM), + "polygon_x": (1, C.NUM_POLYGONS, T.MIXER_T["polygon"], T.MIXER_CIN["polygon"]), + "line_string_x": (1, C.NUM_LINE_STRINGS, T.MIXER_T["line_string"], T.MIXER_CIN["line_string"]), + "goal_x": (1, 1, 1, C.POSE_DIM), + "ego_shape_x": (1, 1, 1, C.EGO_SHAPE_DIM), + "turn_x": (1, 1, 1, C.TURN_INDICATOR_HISTORY), + "token_valid": (1, 1, T.TOKENS, 1), + "pos_aug": (1, 1, T.TOKENS, T.POS_AUG_DIM), + "fusion_key_row": (1, 1, 1, T.TOKENS), + "agent_key_row": (1, 1, 1, T.AGENTS), + "cs": (1, 1, T.AGENTS, T.STATE_COLS), + "y0": (1, 1, T.AGENTS, T.STATE_COLS), +} +BF16_INPUTS = ("fusion_key_row", "agent_key_row") + + +def decoder_state(x: np.ndarray, current_states: Optional[np.ndarray] = None) -> np.ndarray: + """A ``[321, 81, 4]`` state -> ``[1, 1, 352, 324]`` float32 (extra rows zero). With ``current_states`` the t = 0 + columns are replaced by them (the ``cs`` input); with None they are zeroed (the ``y`` state).""" + flat = np.asarray(x, f32).reshape(C.MAX_NUM_AGENTS, T.STATE_COLS) + out = pad_rows(flat, T.AGENTS).astype(f32) + out[:, :T.STATE_COLS_T0] = 0.0 + if current_states is not None: + out[:C.MAX_NUM_AGENTS, :T.STATE_COLS_T0] = np.asarray(current_states, f32) + return out[None, None] + + +def plan_inputs(prep: Prepared) -> Dict[str, np.ndarray]: + """``{input name: array}`` for one plan (see :data:`INPUT_SPECS`).""" + f = prep.features + ego_rows, nb_rows = list(T.ISLAND_ROWS["ego"]), list(T.ISLAND_ROWS["neighbor"]) + tv = pad_rows(np.asarray(f.token_valid, bool), T.TOKENS) + nb_type = np.asarray(f.neighbor_type, f32) + out = { + "ego_x": np.asarray(f.ego, f32)[ego_rows][None, None], + "neighbor_x": np.asarray(f.neighbor, f32)[:, nb_rows][None], + "neighbor_aux": np.concatenate([nb_type, np.ones((nb_type.shape[0], 1), f32)], 1)[None, None], + "static_x": np.asarray(f.static, f32)[None, None], + "lane_x": np.asarray(f.lane, f32)[None], + "lane_aux": lane_aux_features(f.lane_speed, f.lane_has_speed, f.lane_attr)[None, None], + "route_x": np.asarray(f.route, f32)[None], + "route_aux": lane_aux_features(f.route_speed, f.route_has_speed, f.route_attr)[None, None], + "polygon_x": np.asarray(f.polygon, f32)[None], + "line_string_x": np.asarray(f.line_string, f32)[None], + "goal_x": np.asarray(f.goal, f32).reshape(1, 1, 1, -1), + "ego_shape_x": np.asarray(f.ego_shape, f32).reshape(1, 1, 1, -1), + "turn_x": np.asarray(f.turn, f32).reshape(1, 1, 1, -1), + "token_valid": tv.astype(f32).reshape(1, 1, -1, 1), + "pos_aug": np.concatenate([pad_rows(np.asarray(f.pos, f32), T.TOKENS), tv.astype(f32)[:, None]], + 1)[None, None], + "fusion_key_row": key_bias_row(pad_rows(np.asarray(f.key_valid, bool), T.TOKENS)), + "agent_key_row": key_bias_row(pad_rows(np.asarray(prep.decoder.agent_valid, bool), T.AGENTS)), + "cs": decoder_state(np.zeros((C.MAX_NUM_AGENTS, C.OUTPUT_T + 1, C.POSE_DIM), f32), + prep.decoder.current_states), + "y0": decoder_state(prep.x_T), + } + for name, arr in out.items(): + if arr.shape != INPUT_SPECS[name]: + raise AssertionError(f"input {name}: shape {arr.shape}, expected {INPUT_SPECS[name]}") + return out + + +def warmup_inputs() -> Dict[str, np.ndarray]: + """Valid-looking initial contents (every token / agent a valid key: no fully masked softmax row).""" + out = {name: np.zeros(shape, f32) for name, shape in INPUT_SPECS.items()} + out["fusion_key_row"] = key_bias_row(np.arange(T.TOKENS) < T.TOKENS_REAL) + out["agent_key_row"] = key_bias_row(np.arange(T.AGENTS) < T.AGENTS_REAL) + out["token_valid"][..., :T.TOKENS_REAL, 0] = 1.0 + return out diff --git a/code/tt_diffusion_planner/tt/kernels/README.md b/code/tt_diffusion_planner/tt/kernels/README.md new file mode 100644 index 0000000000000000000000000000000000000000..2f43efab1d4f016362bb6a45534d929d2f4b049c --- /dev/null +++ b/code/tt_diffusion_planner/tt/kernels/README.md @@ -0,0 +1,10 @@ +# Custom kernels of diffusion-planner-p150 + +None for the functional port: PLAN.md section 2.12 maps every op of the v5.0 graphs to stock ttnn (matmul / linear, +layer_norm, gelu, SDPA through the shared C20 wrapper, eltwise). Fused kernels (entity-parallel MLP-Mixer, fusion +block, DiT step + solver update) and the single-megakernel attempt (decision D19) are optimization-phase work; when +they land, their `.cpp` sources go here (package data in the repo-root `pyproject.toml`, asserted by a `verify:` line of +`tt-model.yaml`), each with a numpy / torch oracle test and the hang protocol of PLAN.md section 4.4. + +| file | op | used by | status | +|---|---|---|---| diff --git a/code/tt_diffusion_planner/tt/layers.py b/code/tt_diffusion_planner/tt/layers.py new file mode 100644 index 0000000000000000000000000000000000000000..da5e0e0ea8d3ac39e85577519beccc2509824310 --- /dev/null +++ b/code/tt_diffusion_planner/tt/layers.py @@ -0,0 +1,221 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Device building blocks of the planner graph: weights uploaded once with the module's precision, ops that always +carry an explicit compute config (``ttaw.precision``) and LayerNorm epsilon 1e-5 (ttnn's default is 1e-12). + +``Build`` resolves a module name against the precision policy (``tt/config.py`` :data:`DEFAULT_PRECISION` + +``DIFFUSION_PLANNER_PRECISION``): ``w=`` is the weight dtype, ``a=`` the module's residual-stream dtype; hidden MLP +activations are bf16 (``HIDDEN``) unless a module says otherwise. +""" +from __future__ import annotations + +import fnmatch +from typing import Any, Dict, Optional, Sequence + +import numpy as np + +from ..reference import config as C +from ..ttaw.precision import Precision, PrecisionPolicy +from ..ttaw.tensors import to_device, ttnn_dtype +from . import config as T +from .params import Lin, Norm + +__all__ = ["Build", "Linear", "SplitLinear", "make_linear", "LayerNorm", "Const", "ATTN", "policy", + "layer_norm_fp32"] + +ATTN = "bfloat16" # dtype of the Q / K / V projections (SDPA takes bf16) + + +def policy(env: Optional[Dict[str, str]] = None, spec: Optional[str] = None) -> PrecisionPolicy: + """The planner's precision policy: :data:`tt.config.DEFAULT_PRECISION`, then ``DIFFUSION_PLANNER_PRECISION`` (or + ``spec``) rules on top.""" + pol = PrecisionPolicy(dict(T.DEFAULT_PRECISION), default="HiFi4+fp32:w=bf16:a=fp32").with_env( + "DIFFUSION_PLANNER", env) + return pol.override(spec) if spec else pol + + +class Build: + """Device + precision policy + graph options shared by the modules while they upload their weights. + + ``ln_fp32`` / ``hidden_fp32``: module globs (``tt.config.KNOBS`` ``LN_FP32`` / ``HIDDEN_FP32`` by default) whose + LayerNorms use :func:`layer_norm_fp32` / whose hidden MLP activations are fp32.""" + + def __init__(self, device: Any, pol: Optional[PrecisionPolicy] = None, *, ln_fp32: Optional[Sequence[str]] = None, + hidden_fp32: Optional[Sequence[str]] = None, split: Optional[Sequence[str]] = None, + attn_fp32_acc: Optional[Sequence[str]] = None, attn_matmul: Optional[Sequence[str]] = None): + self.device = device + self.policy = pol or policy() + knobs = T.KNOBS.read() + self.ln_fp32 = tuple(T.globs(knobs.LN_FP32) if ln_fp32 is None else ln_fp32) + self.hidden_fp32 = tuple(T.globs(knobs.HIDDEN_FP32) if hidden_fp32 is None else hidden_fp32) + self.split = tuple(T.globs(knobs.SPLIT_MATMUL) if split is None else split) + self.attn_fp32 = tuple(T.globs(knobs.ATTN_FP32_ACC) if attn_fp32_acc is None else attn_fp32_acc) + self.attn_mm = tuple(T.globs(knobs.ATTN_MATMUL) if attn_matmul is None else attn_matmul) + self.uploaded_bytes = 0 + + @staticmethod + def _match(module: str, patterns: Sequence[str]) -> bool: + return any(fnmatch.fnmatchcase(module, p) for p in patterns) + + def ln_mode(self, module: str) -> str: + return "fp32" if self._match(module, self.ln_fp32) else "device" + + def hidden(self, module: str) -> str: + """dtype of the hidden MLP activations of ``module``: fp32 for the ``HIDDEN_FP32`` globs and for split-matmul + modules (a bf16 hidden would undo the split), else bf16.""" + return "float32" if self._match(module, self.hidden_fp32 + self.split) else "bfloat16" + + def split_matmul(self, module: str) -> bool: + """``module``'s fp32 matmuls use :class:`SplitLinear` (~1e-5 instead of the TF32-like ~1e-3).""" + return self._match(module, self.split) + + def attn_fp32_acc(self, module: str) -> bool: + return self._match(module, self.attn_fp32) + + def attn_matmul(self, module: str) -> bool: + """``module``'s attention runs as fp32 matmuls + softmax (C20 ``attention_matmul``).""" + return self._match(module, self.attn_mm) + + def options(self) -> Dict[str, Any]: + return {"ln_fp32": list(self.ln_fp32), "hidden_fp32": list(self.hidden_fp32), "split": list(self.split), + "attn_fp32_acc": list(self.attn_fp32), "attn_matmul": list(self.attn_mm)} + + def prec(self, module: str) -> Precision: + return self.policy.resolve(module) + + def cfg(self, module: str): + return self.policy.compute_kernel_config(module) + + def stream(self, module: str) -> str: + """The residual-stream dtype of ``module``.""" + return self.prec(module).activations + + def upload(self, array: np.ndarray, dtype: str, *, shape4: bool = True): + """Host float array -> DRAM TILE device tensor (rank padded to 4 with leading 1s).""" + a = np.ascontiguousarray(np.asarray(array, np.float32)) + if shape4 and a.ndim < 4: + a = a.reshape((1,) * (4 - a.ndim) + a.shape) + self.uploaded_bytes += a.size * (4 if dtype in ("float32", "fp32") else 2) + return to_device(a, self.device, dtype) + + +class Linear: + """``y = x @ w (+ b)`` with fused ``activation`` ("gelu" = exact erf, "gelu_tanh"); output dtype ``out``.""" + + def __init__(self, build: Build, lin: Lin, module: str, *, out: str, activation: Optional[str] = None, + w_dtype: Optional[str] = None): + p = build.prec(module) + wd = w_dtype or p.weights + self.module, self.activation, self.out = module, activation, ttnn_dtype(out) + self.w = build.upload(lin.w, wd) + self.b = None if lin.b is None else build.upload(np.asarray(lin.b).reshape(1, -1), wd) + self.cfg = build.cfg(module) + self.shape = tuple(lin.w.shape) + + def __call__(self, x): + import ttnn + + return ttnn.linear(x, self.w, bias=self.b, activation=self.activation, dtype=self.out, + compute_kernel_config=self.cfg) + + +def layer_norm_fp32(x, gamma=None, beta=None, *, eps: float = C.LN_EPS): + """LayerNorm over the last dim as fp32 element-wise / reduction ops: ``(x - mean) * rsqrt(var + eps) * gamma + + beta`` (7-8 programs). The fused ``ttnn.layer_norm`` is ~2.5e-3 relative even for fp32 input (probe P10), which + swamps the small per-entity deviations of the offset-dominated mixer rows.""" + import ttnn + + if x.dtype != ttnn.float32: + x = ttnn.typecast(x, ttnn.float32) + xc = ttnn.subtract(x, ttnn.mean(x, dim=-1, keepdim=True)) + var = ttnn.mean(ttnn.multiply(xc, xc), dim=-1, keepdim=True) + y = ttnn.multiply(xc, ttnn.rsqrt(ttnn.add(var, float(eps)), fast_and_approximate_mode=False)) + if gamma is not None: + y = ttnn.multiply(y, gamma) + if beta is not None: + y = ttnn.add(y, beta) + return y + + +class SplitLinear: + """``y = x @ w (+ b)`` to ~1e-5 relative from three device matmuls on split operands (the "bf16x3" scheme): + ``x_hi @ w_hi + x_hi @ w_lo + x_lo @ w_hi`` with ``*_hi`` the bf16 roundings (bf16 x bf16 products are exact + under fp32 accumulation) and ``*_lo = * - *_hi`` the fp32 remainders (exact by Sterbenz), whose TF32-like operand + truncation (probe P12) then costs ~2^-18 of the full value; ``x_lo @ w_lo`` (~2^-16) is dropped. Output fp32, + activation applied to the sum. 8-9 programs instead of 1: for the small pre-projection island only.""" + + def __init__(self, build: Build, lin: Lin, module: str, *, activation: Optional[str] = None, + out: str = "float32"): + from ..ttaw.tensors import round_to_bf16 + + if activation not in (None, "gelu", "gelu_tanh"): + raise ValueError(f"SplitLinear: unsupported activation {activation!r}") + w = np.asarray(lin.w, np.float32) + w_hi = round_to_bf16(w) + self.w_hi = build.upload(w_hi, "bfloat16") + self.w_lo = build.upload((w - w_hi).astype(np.float32), "float32") + self.b = None if lin.b is None else build.upload(np.asarray(lin.b, np.float32).reshape(1, -1), "float32") + self.cfg = build.cfg(module) + self.activation = activation + self.out = ttnn_dtype(out) + self.shape = tuple(w.shape) + + def __call__(self, x): + import ttnn + + f32 = ttnn.float32 + if x.dtype != f32: # a bf16 input is its own hi part: two matmuls + y = ttnn.matmul(x, self.w_hi, dtype=f32, compute_kernel_config=self.cfg) + y = ttnn.add(y, ttnn.linear(x, self.w_lo, bias=self.b, dtype=f32, compute_kernel_config=self.cfg)) + else: + x_hi = ttnn.typecast(x, ttnn.bfloat16) + x_lo = ttnn.subtract(x, ttnn.typecast(x_hi, f32)) + y = ttnn.matmul(x_hi, self.w_hi, dtype=f32, compute_kernel_config=self.cfg) + y = ttnn.add(y, ttnn.matmul(x_hi, self.w_lo, dtype=f32, compute_kernel_config=self.cfg)) + y = ttnn.add(y, ttnn.linear(x_lo, self.w_hi, bias=self.b, dtype=f32, compute_kernel_config=self.cfg)) + if self.activation == "gelu": + y = ttnn.gelu(y, fast_and_approximate_mode=False) + elif self.activation == "gelu_tanh": + y = ttnn.gelu(y, variant=ttnn.GeluVariant.Tanh) + if self.out != f32: + y = ttnn.typecast(y, self.out) + return y + + +def make_linear(build: Build, lin: Lin, module: str, *, out: str, activation: Optional[str] = None): + """:class:`SplitLinear` when ``module`` is in the build's ``SPLIT_MATMUL`` globs, else :class:`Linear`.""" + if build.split_matmul(module): + return SplitLinear(build, lin, module, activation=activation, out=out) + return Linear(build, lin, module, out=out, activation=activation) + + +class LayerNorm: + """LayerNorm with ``epsilon=1e-5`` and fp32 affine rows (``[1, 1, 1, W]`` TILE): ``ttnn.layer_norm``, or + :func:`layer_norm_fp32` when the module is in the build's ``ln_fp32`` globs. ``gamma`` / ``beta`` may be given + as device tensors (the per-step folded adaLN rows of the decoder).""" + + def __init__(self, build: Build, nrm: Optional[Norm], module: str): + self.cfg = build.cfg(module) + self.mode = build.ln_mode(module) + self.g = self.b = None + if nrm is not None: + self.g = build.upload(np.asarray(nrm.gamma).reshape(1, -1), "float32") + self.b = build.upload(np.asarray(nrm.beta).reshape(1, -1), "float32") + + def __call__(self, x, gamma=None, beta=None): + import ttnn + + g = self.g if gamma is None else gamma + b = self.b if beta is None else beta + if self.mode == "fp32": + return layer_norm_fp32(x, g, b) + return ttnn.layer_norm(x, epsilon=C.LN_EPS, weight=g, bias=b, compute_kernel_config=self.cfg) + + +class Const: + """A constant device tensor (upload once).""" + + def __init__(self, build: Build, array: np.ndarray, dtype: str): + self.t = build.upload(array, dtype) + + def __call__(self): + return self.t diff --git a/code/tt_diffusion_planner/tt/model.py b/code/tt_diffusion_planner/tt/model.py new file mode 100644 index 0000000000000000000000000000000000000000..c1b71da245570d8f560307dba4851547b93e41b5 --- /dev/null +++ b/code/tt_diffusion_planner/tt/model.py @@ -0,0 +1,186 @@ +# SPDX-License-Identifier: Apache-2.0 +"""``TtDiffusionPlanner``: the whole plan of Diffusion Planner v5.0 in one trace on a Blackhole p150. + +Variant ``plan`` (``ttaw.trace.TraceRunner``; inputs :data:`tt.inputs.INPUT_SPECS`, written every plan): +encoder + fusion (:mod:`.encoder`) -> cross K / V of the 3 DiT blocks hoisted once -> 11 x (DiT evaluation with +the per-step folded adaLN rows + fp32 DPM-Solver++(2M) update + prefix constraint) (:mod:`.decoder`) -> turn head -> +one packed readback (``final_x0`` ``[352, 324]`` fp32, ``logit`` ``[5]``, ``ego_steps`` ``[11, 324]``). No host +fallback inside the plan, so there are no trace segments. + +Debug variants (``debug=True``; tests only, also replayed from traces): + +- ``encoder_taps``: the encoder of ``plan`` returning every encoder tap of ``reference.model.TAP_NAMES``; +- ``decode_once``: one decoder evaluation on teacher-forced inputs (``dbg_x`` ``[352, 324]``, ``dbg_enc`` + ``[576, 256]``: the reference encoding + 12 zero pad-token rows) with the per-step rows as inputs + (``dbg..``, ``dbg.final.g`` / ``.b``), so one trace serves all 11 evaluation times. + +Everything here runs under the model lock of ``api.DiffusionPlanner`` (one chip, batch 1). +""" +from __future__ import annotations + +import time +from typing import Any, Dict, List, Optional, Sequence + +import numpy as np + +from ..reference import config as C +from ..ttaw.ops import attention as A +from ..ttaw.trace import TraceRunner, pack_outputs +from . import config as T +from . import inputs as I +from . import params as P +from .decoder import STEP_KEYS, TtDecoder, TtTurnHead +from .encoder import TtEncoder +from .layers import Build, policy + +__all__ = ["TtDiffusionPlanner"] + + +class TtDiffusionPlanner: + """Weights on the device, the ``TraceRunner`` with its persistent inputs and variants, and the host glue. + + ``weights``: ``reference.weights.PlannerWeights``; ``precision``: extra policy rules (``"dec.*=HiFi2+fp32"``) + on top of ``DIFFUSION_PLANNER_PRECISION``; ``ln_fp32`` / ``hidden_fp32`` / ``split`` / ``attn_fp32_acc`` / + ``attn_matmul``: module globs overriding the knobs of the same names (``tt/config.py``); ``debug``: also register + ``encoder_taps`` / ``decode_once``.""" + + def __init__(self, device: Any, weights: Any, *, debug: bool = False, precision: Optional[str] = None, + ln_fp32: Optional[Sequence[str]] = None, hidden_fp32: Optional[Sequence[str]] = None, + split: Optional[Sequence[str]] = None, attn_fp32_acc: Optional[Sequence[str]] = None, + attn_matmul: Optional[Sequence[str]] = None, steps: int = C.DPM_SOLVER_STEPS, + num_command_queues: Optional[int] = None): + import ttnn + + t0 = time.perf_counter() + self.device = device + self.build = Build(device, policy(spec=precision), ln_fp32=ln_fp32, hidden_fp32=hidden_fp32, split=split, + attn_fp32_acc=attn_fp32_acc, attn_matmul=attn_matmul) + p = weights.params + self.tables = P.step_tables(p, steps) + self.encoder = TtEncoder(self.build, p) + self.decoder = TtDecoder(self.build, p, self.tables) + self.turn = TtTurnHead(self.build, p) + self.runner = TraceRunner(device, num_command_queues=num_command_queues, name="diffusion-planner") + warm = I.warmup_inputs() + for name, shape in I.INPUT_SPECS.items(): + dtype = "bfloat16" if name in I.BF16_INPUTS else "float32" + self.runner.add_input(name, init=warm[name], dtype=dtype, layout=ttnn.TILE_LAYOUT) + self.runner.add_variant("plan", self._plan) + self.debug = bool(debug) + if self.debug: + self.runner.add_input("dbg_x", init=np.zeros((1, 1, T.AGENTS, T.STATE_COLS), np.float32), + dtype="float32", layout=ttnn.TILE_LAYOUT) + self.runner.add_input("dbg_enc", init=np.zeros((1, 1, T.TOKENS, C.HIDDEN_DIM), np.float32), + dtype=self.encoder.fstream, layout=ttnn.TILE_LAYOUT) + for name in self._dbg_row_names(): + self.runner.add_input(name, init=np.zeros((1, 1, 1, C.HIDDEN_DIM), np.float32), dtype="float32", + layout=ttnn.TILE_LAYOUT) + self.runner.add_variant("encoder_taps", self._encoder_taps) + self.runner.add_variant("decode_once", self._decode_once) + self.build_ms = (time.perf_counter() - t0) * 1e3 + + # ---- traced functions -------------------------------------------------------------------------------------- + def _plan(self, ctx): + enc = self.encoder.forward(ctx) + kv = self.decoder.cross_kv(enc) + self_mask = A.expand_key_bias(ctx["agent_key_row"], T.AGENTS) + final, ego = self.decoder.solve(ctx["y0"], ctx["cs"], kv, self_mask) + logit = self.turn(final, enc) + return pack_outputs({"final_x0": final, "logit": logit, "ego_steps": ego}) + + def _encoder_taps(self, ctx): + taps: Dict[str, Any] = {} + self.encoder.forward(ctx, taps) + return taps + + @staticmethod + def _dbg_row_names() -> List[str]: + return [f"dbg.{i}.{k}" for i in range(C.DIT_DEPTH) for k in STEP_KEYS] + ["dbg.final.g", "dbg.final.b"] + + def _decode_once(self, ctx): + rows = {"blocks": [{k: ctx[f"dbg.{i}.{k}"] for k in STEP_KEYS} for i in range(C.DIT_DEPTH)], + "g": ctx["dbg.final.g"], "b": ctx["dbg.final.b"]} + kv = self.decoder.cross_kv(ctx["dbg_enc"]) + self_mask = A.expand_key_bias(ctx["agent_key_row"], T.AGENTS) + return self.decoder.evaluate(ctx["dbg_x"], rows, kv, self_mask) + + # ---- host side ------------------------------------------------------------------------------------------- + def capture(self) -> None: + self.runner.capture() + + def _run(self, variant: str, inputs: Dict[str, Any], eager: bool): + return self.runner.run_eager(variant, inputs=inputs) if eager else self.runner(variant, inputs=inputs) + + def forward(self, prepared: Any, *, ego_steps: bool = True, eager: bool = False) -> Dict[str, Any]: + """One plan: upload, replay, read (``eager=True``: the same graph without the trace, for bring-up and + replay-vs-eager checks). ``final_x0`` ``[321, 81, 4]`` (normalised, prefix-constrained), ``logit`` ``[5]``, + ``denoising_steps``: the 11 iterates' ego rows as ``[1, 81, 4]`` arrays.""" + out = self._run("plan", I.plan_inputs(prepared), eager) + final = out["final_x0"].reshape(T.AGENTS, T.STATE_COLS)[:C.MAX_NUM_AGENTS] + res = {"final_x0": final.reshape(C.MAX_NUM_AGENTS, C.OUTPUT_T + 1, C.POSE_DIM).astype(np.float32), + "logit": out["logit"].reshape(-1)[:C.TURN_INDICATOR_OUTPUT_DIM].astype(np.float32)} + if ego_steps: + steps = out["ego_steps"].reshape(-1, T.STATE_COLS) + res["denoising_steps"] = [s.reshape(1, C.OUTPUT_T + 1, C.POSE_DIM).astype(np.float32) for s in steps] + return res + + def encoder_taps(self, prepared: Any, *, eager: bool = False) -> Dict[str, np.ndarray]: + """Replay of ``encoder_taps``: ``{tap: array}`` with the reference's shapes (``[E, 256]`` category rows, + ``[E, 64, 128]`` mixer taps, ``[564, 256]`` token / fusion / encoding rows).""" + if not self.debug: + raise RuntimeError("encoder_taps needs TtDiffusionPlanner(debug=True)") + raw = self._run("encoder_taps", I.plan_inputs(prepared), eager) + out = {} + for name, a in raw.items(): + a = np.asarray(a, np.float32) + if name.endswith((".pre", ".mixer")): + out[name] = a.reshape(-1, C.MIXER_TOKENS, C.MIXER_CHANNELS) + elif name in ("enc.tokens", "enc.encoding") or name.startswith("enc.fusion."): + out[name] = a.reshape(-1, C.HIDDEN_DIM)[:T.TOKENS_REAL] + else: + out[name] = a.reshape(-1, C.HIDDEN_DIM) + return out + + def decode_once(self, prepared: Any, x: np.ndarray, t: float, *, encoding: np.ndarray, + eager: bool = False) -> np.ndarray: + """Replay of ``decode_once`` at evaluation time ``t`` (one of ``tables.eval_times``) on the teacher-forced + ``x`` ``[321, 81, 4]`` (prefix-constrained) and ``encoding`` ``[564, 256]`` -> ``[321, 81, 4]`` (the t = 0 + slot is 0: the masked projection).""" + if not self.debug: + raise RuntimeError("decode_once needs TtDiffusionPlanner(debug=True)") + k = int(np.argmin([abs(float(t) - e) for e in self.tables.eval_times])) + if abs(float(t) - self.tables.eval_times[k]) > 1e-6: + raise ValueError(f"t={t} is not an evaluation time of the solver plan {self.tables.eval_times}") + base = I.plan_inputs(prepared) + enc = P.pad_rows(np.asarray(encoding, np.float32).reshape(T.TOKENS_REAL, C.HIDDEN_DIM), T.TOKENS) + inputs = {"agent_key_row": base["agent_key_row"], "dbg_x": I.decoder_state(x, x[:, 0]), + "dbg_enc": enc.reshape(1, 1, T.TOKENS, C.HIDDEN_DIM)} + for i, blk in enumerate(self.tables.blocks[k]): + for key in STEP_KEYS: + inputs[f"dbg.{i}.{key}"] = np.asarray(blk[key], np.float32).reshape(1, 1, 1, -1) + inputs["dbg.final.g"] = np.asarray(self.tables.final[k]["g"], np.float32).reshape(1, 1, 1, -1) + inputs["dbg.final.b"] = np.asarray(self.tables.final[k]["b"], np.float32).reshape(1, 1, 1, -1) + out = self._run("decode_once", inputs, eager) + flat = np.asarray(out, np.float32).reshape(T.AGENTS, T.STATE_COLS)[:C.MAX_NUM_AGENTS] + return flat.reshape(C.MAX_NUM_AGENTS, C.OUTPUT_T + 1, C.POSE_DIM) + + def trace_buffers_mb(self) -> Optional[float]: + """DRAM held by the captured traces' command buffers: the allocated bytes of the device's TRACE region + (``ttnn.get_memory_view``; all banks), None when the view is unavailable.""" + import ttnn + + try: + view = ttnn.get_memory_view(self.device, ttnn.BufferType.TRACE) + return round(view.num_banks * view.total_bytes_allocated_per_bank / 2 ** 20, 2) + except (AttributeError, RuntimeError, TypeError): + return None + + def describe(self) -> Dict[str, Any]: + return {"trace": self.runner.describe(), "trace_buffers_mb": self.trace_buffers_mb(), + "precision": self.build.policy.describe(), + "options": self.build.options(), + "uploaded_mb": round(self.build.uploaded_bytes / 2 ** 20, 2), "build_ms": round(self.build_ms, 1), + "tokens": T.TOKENS, "agents": T.AGENTS, "nfe": self.tables.nfe} + + def release(self) -> None: + self.runner.release() diff --git a/code/tt_diffusion_planner/tt/params.py b/code/tt_diffusion_planner/tt/params.py new file mode 100644 index 0000000000000000000000000000000000000000..9ad1fc682ca79f0a4f2e1e0092443f0a9044cf9a --- /dev/null +++ b/code/tt_diffusion_planner/tt/params.py @@ -0,0 +1,222 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Host-side (numpy) preparation of every tensor the ttnn graph holds: the canonical weights of +``reference.weights`` plus the exact rewrites of the port, folded in float64 and rounded to float32 once. + +No ttnn here (host-testable); the modules of :mod:`.encoder` / :mod:`.decoder` upload what they need with their +precision. Rewrites (each exact in real arithmetic; ``tests/test_tt_params_host.py`` proves them against the CPU +reference): + +- **pad-relative island** (ego / neighbour ``channel_pre`` + ``token_pre``): ``reference.rewrites.island_constants``; +- **neighbour type embedding** ``n + type @ W_t + b_t`` = ``n + [type, 1] @ [W_t; b_t]`` (:func:`neighbor_aux`); +- **lane / route speed + attribute embeddings**: ``where(has, speed * w_s + b_s, unk) + attr @ W_a + b_a`` with + ``has`` in {0, 1} = ``[speed * has, has, 1 - has, attr, 1] @ [w_s; b_s; unk; W_a; b_a]`` (:func:`lane_aux`); +- **masked positional embedding** ``valid * (pos @ W + b)`` with invalid pos rows already zero = + ``[pos, valid] @ [W; b]`` (:func:`pos_aug`); +- **agent embedding** folded into the pre-projection bias rows: ``preproj.fc2(h) + emb[ego / neighbour]`` = + ``h @ W2 + rows`` with ``rows[0] = b2 + emb[0]``, ``rows[i > 0] = b2 + emb[1]`` (:func:`agent_rows`); +- **masked last projection**: the t = 0 output columns (0..3) of ``final_layer.proj.4`` zeroed, so the model output + is ``m * mask0`` (the prefix constraint overwrites that slot anyway; ``tt/config.py`` solver state); +- **turn head** ``W [272 -> 5]`` over ``(final_x0[0, 1::10, :2], mean(encoding))`` = ``x0_row0 @ W_sel + + sum_tokens(encoding) @ (W_pool / 564) + b`` with ``W_sel`` the 16 used rows scattered to ``[324, 5]`` + (:func:`turn_weights`); +- **per-step adaLN tables** folded into the LayerNorm affine (``reference.rewrites.adaln_tables``) and the solver + update ``x' = a x - b m0 - c (m0 - m1) / r0`` as ``y' = A y - B m0 + Cm m1`` with ``A = a``, ``B = b + c / r0``, + ``Cm = c / r0`` (:func:`solver_coefficients`; float64, rounded once). +""" +from __future__ import annotations + +from dataclasses import dataclass +from typing import Dict, List, Mapping, Optional, Tuple + +import numpy as np + +from ..host.solver import SolverPlan, solver_plan +from ..reference import config as C +from ..reference import rewrites as R +from . import config as T + +__all__ = ["Lin", "Norm", "linear", "norm", "mixer_module", "neighbor_aux", "lane_aux", "lane_aux_features", + "pos_aug", "agent_rows", "final_projection_masked", "turn_weights", "solver_coefficients", "StepTables", + "step_tables", "island", "small_module", "pad_rows", "MIXER_MODULES", "SMALL_MODULES"] + +MIXER_MODULES = {"ego": "encoder.ego_encoder", "neighbor": "encoder.neighbor_encoder", + "lane": "encoder.lane_encoder", "route": "encoder.route_encoder", + "polygon": "encoder.polygon_encoder", "line_string": "encoder.line_string_encoder"} +SMALL_MODULES = {"goal": "encoder.goal_pose_encoder", "ego_shape": "encoder.ego_shape_encoder", + "turn": "encoder.turn_indicator_encoder"} + +f32 = np.float32 + + +@dataclass +class Lin: + """``y = x @ w + b``; ``w`` [in, out] float32, ``b`` [out] or None.""" + + w: np.ndarray + b: Optional[np.ndarray] + + +@dataclass +class Norm: + gamma: np.ndarray + beta: np.ndarray + + +def _a(p: Mapping[str, np.ndarray], name: str) -> np.ndarray: + return np.ascontiguousarray(np.asarray(p[name], f32)) + + +def linear(p: Mapping[str, np.ndarray], name: str, *, bias: bool = True) -> Lin: + return Lin(_a(p, f"{name}.w"), _a(p, f"{name}.b") if bias else None) + + +def norm(p: Mapping[str, np.ndarray], name: str) -> Norm: + return Norm(_a(p, f"{name}.gamma"), _a(p, f"{name}.beta")) + + +def mixer_module(p: Mapping[str, np.ndarray], cat: str) -> Dict[str, object]: + """Weights of one MLP-Mixer category: ``channel_pre`` / ``token_pre`` (fc1, fc2), 6 blocks (norm1, tokens fc1 / + fc2, norm2, channels fc1 / fc2), ``norm``, ``emb_project`` (fc1, fc2).""" + N = MIXER_MODULES[cat] + out: Dict[str, object] = { + "c1": linear(p, f"{N}.channel_pre_project.fc1"), "c2": linear(p, f"{N}.channel_pre_project.fc2"), + "t1": linear(p, f"{N}.token_pre_project.fc1"), "t2": linear(p, f"{N}.token_pre_project.fc2"), + "norm": norm(p, f"{N}.norm"), "e1": linear(p, f"{N}.emb_project.fc1"), "e2": linear(p, f"{N}.emb_project.fc2"), + "blocks": [{"n1": norm(p, f"{N}.blocks.{i}.norm1"), "tk1": linear(p, f"{N}.blocks.{i}.tokens_mlp.fc1"), + "tk2": linear(p, f"{N}.blocks.{i}.tokens_mlp.fc2"), "n2": norm(p, f"{N}.blocks.{i}.norm2"), + "ch1": linear(p, f"{N}.blocks.{i}.channels_mlp.fc1"), + "ch2": linear(p, f"{N}.blocks.{i}.channels_mlp.fc2")} for i in range(C.MIXER_DEPTH)], + } + return out + + +def island(p: Mapping[str, np.ndarray], cat: str) -> Dict[str, np.ndarray]: + """Pad-relative island of ``cat`` in {"ego", "neighbor"} (``reference.rewrites``, probe P12): ``c1`` (fc1 of + channel_pre, with bias), ``gelu_b1`` [1, 128], ``c2_w`` [128, 128] (no bias: deviation), ``t1_w`` = W_t1 of the + 6 kept rows [6, 64], ``t1_pad`` / ``g_pad`` / ``t2_pad`` [128, 64], ``t2_w`` [64, 64].""" + N = MIXER_MODULES[cat] + cst = R.island_constants(p, cat) + if tuple(cst.valid_rows) != T.ISLAND_ROWS[cat]: + raise AssertionError(f"island rows of {cat} differ from tt.config.ISLAND_ROWS") + return {"c1_w": _a(p, f"{N}.channel_pre_project.fc1.w"), "c1_b": _a(p, f"{N}.channel_pre_project.fc1.b"), + "gelu_b1": cst.gelu_b1.reshape(1, -1), "c2_w": _a(p, f"{N}.channel_pre_project.fc2.w"), + "t1_w": np.ascontiguousarray(cst.w_t1_valid), "t1_pad": cst.t1_pad, "g_pad": cst.g_pad, + "t2_pad": cst.t2_pad, "t2_w": _a(p, f"{N}.token_pre_project.fc2.w")} + + +def neighbor_aux(p: Mapping[str, np.ndarray]) -> np.ndarray: + """``[W_type; b_type]`` [4, 128] for the host's ``[type one-hot, 1]`` columns.""" + N = MIXER_MODULES["neighbor"] + return np.concatenate([_a(p, f"{N}.type_emb.w"), _a(p, f"{N}.type_emb.b")[None]], 0).astype(f32) + + +def lane_aux(p: Mapping[str, np.ndarray], cat: str) -> np.ndarray: + """``[w_s; b_s; unk; W_a; b_a]`` [29, 128] for ``[speed * has, has, 1 - has, attributes (25), 1]``.""" + N = MIXER_MODULES[cat] + rows = [_a(p, f"{N}.speed_limit_emb.w").reshape(1, -1), _a(p, f"{N}.speed_limit_emb.b")[None], + _a(p, f"{N}.unknown_speed_emb").reshape(1, -1), _a(p, f"{N}.attribute_emb.w"), + _a(p, f"{N}.attribute_emb.b")[None]] + out = np.concatenate(rows, 0).astype(f32) + assert out.shape == (T.LANE_AUX_DIM, C.MIXER_CHANNELS), out.shape + return out + + +def lane_aux_features(speed: np.ndarray, has_speed: np.ndarray, attr: np.ndarray) -> np.ndarray: + """Host columns of :func:`lane_aux`: ``[E, 29]`` float32 (``has`` is the node's ``> FLT_EPSILON`` mask).""" + s = np.asarray(speed, f32).reshape(-1, 1) + h = np.asarray(has_speed, bool).reshape(-1, 1).astype(f32) + a = np.asarray(attr, f32).reshape(s.shape[0], -1) + return np.concatenate([s * h, h, f32(1.0) - h, a, np.ones_like(h)], axis=1).astype(f32) + + +def pos_aug(p: Mapping[str, np.ndarray]) -> np.ndarray: + """``[W_pos; b_pos]`` [15, 256] for the host's ``[pos (14), token valid]`` columns.""" + return np.concatenate([_a(p, "encoder.pos_emb.w"), _a(p, "encoder.pos_emb.b")[None]], 0).astype(f32) + + +def agent_rows(p: Mapping[str, np.ndarray], agents: int = T.AGENTS) -> np.ndarray: + """``preproj.fc2`` bias + agent embedding per decoder row [agents, 256] (row 0 ego, others neighbour), folded in + float64 and rounded once.""" + b2 = np.asarray(p["decoder.dit.preproj.fc2.b"], np.float64) + emb = np.asarray(p["decoder.dit.agent_embedding"], np.float64) + rows = np.repeat((b2 + emb[1])[None], agents, axis=0) + rows[0] = b2 + emb[0] + return rows.astype(f32) + + +def final_projection_masked(p: Mapping[str, np.ndarray]) -> Lin: + """``final_layer.proj.4`` (1024 -> 324) with the t = 0 output columns zeroed.""" + lin = linear(p, "decoder.dit.final_layer.proj.4") + w, b = lin.w.copy(), lin.b.copy() + w[:, :T.STATE_COLS_T0] = 0.0 + b[:T.STATE_COLS_T0] = 0.0 + return Lin(w, b) + + +def turn_weights(p: Mapping[str, np.ndarray]) -> Dict[str, np.ndarray]: + """``w_sel`` [324, 5] (the 16 used rows of W scattered to the ``(t, d)`` columns of a state row), ``w_pool`` + [256, 5] = ``W[16:] / 564`` (float64, rounded once), ``b`` [5].""" + w = np.asarray(p["decoder.turn_indicator_predictor.w"], np.float64) # [272, 5] + w_sel = np.zeros((T.STATE_COLS, w.shape[1]), np.float64) + for i, t in enumerate(C.TURN_HEAD_STEPS): + for d in range(2): + w_sel[t * C.POSE_DIM + d] = w[2 * i + d] + n_sel = 2 * len(C.TURN_HEAD_STEPS) + return {"w_sel": w_sel.astype(f32), "w_pool": (w[n_sel:] / T.TOKENS_REAL).astype(f32), + "b": np.asarray(p["decoder.turn_indicator_predictor.b"], f32)} + + +def solver_coefficients(plan: SolverPlan) -> List[Tuple[float, float, float]]: + """``(A, B, Cm)`` per update (float32 values) of ``y' = A y - B m0 + Cm m1`` (``Cm = 0`` for the first-order + update).""" + out = [] + for u in plan.updates: + a, b, c, r0 = (float(v) for v in (u.a, u.b, u.c, u.r0)) + cm = c / r0 if u.order > 1 else 0.0 + out.append((float(f32(a)), float(f32(b + cm)), float(f32(cm)))) + return out + + +@dataclass +class StepTables: + """Per evaluation ``k``: per DiT block the folded ``norm1`` / ``norm2`` affine rows and the gates, and the folded + ``norm_final`` affine; ``solver`` = :func:`solver_coefficients` (one fewer than evaluations).""" + + eval_times: Tuple[float, ...] + blocks: List[List[Dict[str, np.ndarray]]] # [k][i] -> {n1_g, n1_b, gate_msa, n2_g, n2_b, gate_mlp} + final: List[Dict[str, np.ndarray]] # [k] -> {g, b} + solver: List[Tuple[float, float, float]] + plan: SolverPlan + + @property + def nfe(self) -> int: + return len(self.eval_times) + + +def step_tables(p: Mapping[str, np.ndarray], steps: int = C.DPM_SOLVER_STEPS) -> StepTables: + plan = solver_plan(steps) + tab = R.adaln_tables(p, plan.eval_times) + blocks = [[{"n1_g": blk["norm1_gamma"][k], "n1_b": blk["norm1_beta"][k], "gate_msa": blk["gate_msa"][k], + "n2_g": blk["norm2_gamma"][k], "n2_b": blk["norm2_beta"][k], "gate_mlp": blk["gate_mlp"][k]} + for blk in tab.blocks] for k in range(len(plan.eval_times))] + final = [{"g": tab.final_gamma[k], "b": tab.final_beta[k]} for k in range(len(plan.eval_times))] + return StepTables(tuple(plan.eval_times), blocks, final, solver_coefficients(plan), plan) + + +def small_module(p: Mapping[str, np.ndarray], cat: str) -> Dict[str, object]: + """goal / ego-shape / turn encoders: channel MLP (fc1, fc2), ``norm``, ``emb_project`` (fc1, fc2).""" + N = SMALL_MODULES[cat] + return {"c1": linear(p, f"{N}.channel_pre_project.fc1"), "c2": linear(p, f"{N}.channel_pre_project.fc2"), + "norm": norm(p, f"{N}.norm"), "e1": linear(p, f"{N}.emb_project.fc1"), + "e2": linear(p, f"{N}.emb_project.fc2")} + + +def pad_rows(a: np.ndarray, rows: int) -> np.ndarray: + """Zero-pad the first axis to ``rows``.""" + a = np.asarray(a) + if a.shape[0] > rows: + raise ValueError(f"{a.shape[0]} rows exceed {rows}") + out = np.zeros((rows,) + a.shape[1:], a.dtype) + out[:a.shape[0]] = a + return out + diff --git a/code/tt_diffusion_planner/ttaw/API.md b/code/tt_diffusion_planner/ttaw/API.md new file mode 100644 index 0000000000000000000000000000000000000000..3ed0782e3e79ed808a1a6cd607cd456d62704b87 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/API.md @@ -0,0 +1,1272 @@ +# ttaw API guide (v0.20.0) + +`ttaw` is the shared code of the 13 Autoware ports to one Blackhole p150b (PLAN.md section 1). Source of truth: +`common/ttaw/`. Every bundle gets a vendored copy at `bundles/-p150/code//ttaw/` and imports it +relatively (`from .ttaw.trace import TraceRunner`). Tests: `common/tests/host` (CPU, fake ttnn) and +`common/tests/device` (p150, through `bin/devrun`). Changes: `common/CHANGELOG.md`. + +Rules that keep one tree valid top-level and vendored: + +- relative imports only; importing any module never imports ttnn / torch / onnx, opens no device, reads no + environment and touches no network (host tests and the image's `verify:` step import everything without a device); +- every optimization knob is read once at model build and has an env A/B switch (`ttaw.knobs`); +- kernel `.cpp` files are package data under `ops/kernels/`, located with `ops.kernel_path(name)`. + +Contents: 1. bundle wiring - 2. device - 3. TraceRunner - 4. tensors - 5. weights - 6. precision and knobs - +7. metrics, goldens, gates - 8. profiling - 9. I/O and outputs - 10. ModelBase - 11. server - 12. vendoring - +13. measured facts and pitfalls - 14. image pre-processing (C14) - 15. LiDAR host pipeline (C10-C12, C15, C16) - +16. CNN ops: conv builders (C17), up-sampling and interpolation (C18) - 17. attention (C20) - +18. LiDAR device modules: gather-form scatter (C19), SECOND + SECONDFPN (C24), CenterHead (C25) - +19. grid_sample helpers (C23) - 20. ResNet builders (C27) - 21. query heads: top-k (C21), heatmap peaks (C22), +TransFusion query head (C26) - 22. pillar feature net and input staging (C24 companion) - 23. segment reductions (K1). + +--------------------------------------------------------------------------------------------------------------------- + +## 1. Wiring a bundle (thin wrappers over ttaw) + +```bash +python common/tools/vendor.py bundles/-p150 # committed ttaw/ (HEAD) -> code//ttaw + VENDORED.json +python common/tools/vendor.py bundles/-p150 --check # drift report (exit 1 on drift; section 12) +``` + +`research/BUNDLE_TEMPLATE` is the canonical wiring (`instantiate_bundle.py --out` renders it and vendors ttaw; the +values file's `INPUT_KIND` picks its flavour: `lidar` as below, `camera`, `multi_camera`, `lidar_camera`, `planner` +with their own `api.py` / `io.py` / stub / smoke / tests, `research/BUNDLE_TEMPLATE/TEMPLATE_NOTES.md` "Flavours"): +`device.py` and `io.py` bind ttaw to the model, `api.py` subclasses `ModelBase`, `server/app.py` calls `create_app`, +and the stdlib `server/client.py` / `server/smoke_test.py` run `ttaw/server/{client,smoke}.py` by path: + +```python +# code//device.py: ttaw.device bound to the port's validated open parameters (the model class uses the same dict) +from .ttaw.device import DeviceConfig, close_device, describe_device # noqa: F401 +DEVICE_DEFAULTS = {"num_command_queues": 1, "l1_small_size": 32768, "trace_region_size": 64 << 20} +def device_config(**overrides): ... # DEVICE_DEFAULTS < _* / TT_DEVICE_ID < non-None overrides +def open_device(device_id=None, *, dispatch=None, **overrides): ... # device_config(...).open() + +# code//io.py: ttaw.io re-exported, decode_points / load_points rebound to the model's layout +from .ttaw.io import * # noqa: F401,F403 +DEFAULT_POINT_FIELDS = ("x", "y", "z", "intensity") + +# code//api.py (section 10 has the full hook contract) +from .ttaw.api_base import ModelBase +class CenterPoint(ModelBase): + MODEL_NAME, ENV_PREFIX = "centerpoint-p150", "CENTERPOINT" + DEVICE_DEFAULTS = device.DEVICE_DEFAULTS + ... + +# code//server/app.py (section 11) +from pathlib import Path +from ..api import CenterPoint +from ..ttaw.server.app import ServerSpec, create_app, parse_mesh_shape # noqa: F401 (tt-model.yaml verify: line) +app = create_app(ServerSpec(model_name=CenterPoint.MODEL_NAME, env_prefix="CENTERPOINT", model_cls=CenterPoint, + task="LiDAR 3D object detection", default_weights=CenterPoint.DEFAULT_REPO, + calib_dir=Path(__file__).resolve().parents[1] / "calib")) +``` + +`pyproject.toml` package data: `".ttaw" = ["API.md", "VENDORED.json"]`, `".ttaw.ops" = ["kernels/*.cpp", +"kernels/*.hpp", "kernels/*.h"]`. The tt-model `extra_code` path `` already ships the sub-package. + +--------------------------------------------------------------------------------------------------------------------- + +## 2. Device (C01, `ttaw.device`) + +```python +from .ttaw.device import DeviceConfig, describe_device, device_session, open_device, compute_grid + +with device_session(dispatch="eth", num_command_queues=2, trace_region_size=64 << 20) as dev: + print(describe_device(dev)) # {'dispatch': 'eth', 'grid': '12x10', 'cores': 120, 'num_command_queues': 2, + # 'arch': 'blackhole', 'fallback': None, 'eth_patch': True, ...} + gx, gy = compute_grid(dev) # never hard-code 12x10: program configs read the grid + +cfg = DeviceConfig.from_env("CENTERPOINT", num_command_queues=1) # _DISPATCH / _NUM_CQS / _L1_SMALL / +dev = cfg.open() # _TRACE_REGION / _WORKER_L1_SIZE, TT_DEVICE_ID +``` + +| name | signature / meaning | +|---|---| +| `open_device` | `(device_id=0, *, dispatch="eth", num_command_queues=1, l1_small_size=32768, trace_region_size=64<<20, worker_l1_size=None, allow_fallback=True)` -> ttnn device. `dispatch`: `eth` (12x10), `worker` (11x10, A/B only), `auto` (ETH if the patch marker is in `tt_metal/impl/dispatch/topology.cpp`, else WORKER + `RuntimeWarning`). A failed ETH open warns and falls back to WORKER; `allow_fallback=False` re-raises (use it in container smoke / CI). Warns if ETH does not give 12x10 | +| `DeviceConfig` | frozen dataclass of the arguments above; `.from_env(prefix, env=None, **defaults)`, `.open()` | +| `describe_device` | `(device, device_id=None)` -> dict: `dispatch` (used), `dispatch_requested`, `fallback`, `num_command_queues`, `grid` "12x10", `grid_x`, `grid_y`, `cores`, `arch`, open sizes, `eth_patch`. `dispatch="unknown"` for devices not opened by ttaw | +| `open_info(device)` | what `open_device` recorded for this device object (`{}` for a device opened elsewhere; the record is tied to the object, never to a reused `id`) | +| `close_device` / `device_session` | close (syncs first; idempotent: a second close of the same device is a no-op) / context manager that always closes | +| `resolve_dispatch(dispatch, *, home=None)`, `eth_dispatch_patch_present(home=None)`, `tt_metal_home()` | patch detection (file read; `TT_METAL_HOME`, `TT_METAL_RUNTIME_ROOT`, or the tree holding `ttnn`, found without importing it) | +| `compute_grid(dev)` / `core_grid(dev)` / `full_core_range_set(dev)` | `(x, y)` / `ttnn.CoreGrid` / one-rectangle `ttnn.CoreRangeSet` of the whole grid | +| constants | `DISPATCH_MODES`, `PATCH_MARKER`, `DEFAULT_L1_SMALL_SIZE`, `DEFAULT_TRACE_REGION_SIZE`, `P150_ETH_GRID` | + +`worker_l1_size` is absolute bytes of allocatable L1 per core; tt-metal's default is computed at open time +(about 1,461,248 B on this tree, TT_PLATFORM.md section 1). Shrinking it grows the kernel-config ring buffer. + +--------------------------------------------------------------------------------------------------------------------- + +## 3. Traces (C02, `ttaw.trace`) + +Every device stage runs inside traces. `TraceRunner` owns the persistent tensors, the warm-up, the captures, the +replays and the readback. + +```python +import ttnn +from .ttaw.trace import TraceRunner, pack_outputs + +runner = TraceRunner(device, num_command_queues=2, name="centerpoint") # CQ count defaults to the open's +runner.add_input("pillars", shape=(1, 1, 40000, 64), dtype="bfloat16") # ROW_MAJOR in DRAM by default +runner.add_input("index", init=warm_idx, dtype="uint32") # give real data if the graph gathers +runner.add_param("score_thr", 0.35) # RT-dev: fp32 [1,1,1,1] TILE +runner.add_state("prev_bev", shape=(1, 1, 22500, 256), dtype="bfloat16", pingpong=True) # temporal models + +def forward(ctx): # ctx[name]: input / param / state read buffer; no host I/O in here + x = model.backbone(ttnn.to_layout(ctx["pillars"], ttnn.TILE_LAYOUT)) + bev = model.temporal(x, ctx["prev_bev"]) + ctx.write_state("prev_bev", bev) # ttnn.copy into the state's write buffer (part of the trace) + heat = ttnn.gt(model.head(bev), ctx["score_thr"]) + return pack_outputs({"heat": heat, "boxes": model.box(bev)}) # one D2H for several outputs + +runner.add_variant("default", forward) # one trace per variant (buckets: "n8k", "n16k", ...) +runner.capture() # warm-up of every variant, then strict captures +out = runner("default", inputs={"pillars": p, "index": idx}, params={"score_thr": 0.4}) # {"heat": ..., "boxes": ...} +runner.release() # or `with TraceRunner(...) as runner:` +``` + +**What it guarantees.** Persistent tensors exist (with defined contents) before the first capture; every variant is +warmed `warmup_runs` times before *any* capture, and so are the runner's own eager programs (the `stage_inputs` +staging copy, the state <-> bank copies) and the callables registered with `add_eager_warmup(fn)`; capture runs with +`device.set_program_cache_misses_allowed(False)` +(a miss raises `... program cache miss occurred, but cache misses are forbidden` naming the op), `end_trace_capture` +runs in `finally` and a failed capture is released, and the program-cache size must not change during capture; +adding a variant after a capture releases all traces, warms the new one and recaptures everything; warm-up writes to +states are undone (`reset_state`) before capture; unchanged params are not re-uploaded. + +**No program may be compiled after the first capture.** A new program's kernel binaries are a DRAM buffer allocated +when it is first enqueued, in the address space of the traces' freed intermediates, so a replay can overwrite them +(tt-metal `tech_reports/AdvancedPerformanceOptimizationsForModels/TraceCorrectness.md`: corruption or a hang). Any +eager op the model runs between replays (host-fallback glue, an eager `ttnn.to_layout`, `tensors.upload_u8`, +`fp32_island`, ...) must therefore run once before the first capture: register it with `add_eager_warmup(fn)`. The +runner refuses its own eager work that compiles after a capture (`stage_fn` copies, `save_state` / `load_state`, +`reset_state` from a device tensor of another spec, `run_eager`) with `RuntimeError: ... compiled a new program after +capture`, and every port runs its device tests once with `TT_METAL_TRACE_ALLOC_TRACKING=1` (the tracker refuses a +replay while such a buffer is alive; ttaw 0.2.0's `stage_inputs` path tripped it: +`logs/ttaw/review_old_staged_tracking_probe.log`). + +**1CQ / 2CQ.** 1CQ: uploads and replays on CQ0. 2CQ: CQ1 waits for the last replay's event, uploads, records; CQ0 +waits, replays, records (the `check_dispatch.py` pattern; the events exist even after a partially failed `capture()` +left some variants runnable). `stage_inputs=True` (2CQ only) is the `tt_cnn` executor +pattern: CQ1 writes DRAM staging copies while the previous replay still runs, an eager `ttnn.copy` (or your +`add_input(..., stage_fn=lambda staging, dst: ...)`, e.g. a reshard into a sharded L1 input) moves them into the +trace inputs on CQ0; that copy is compiled in the first `capture()` (a staging buffer mirrors every +`write_input`). `read(variant, cq_id=1)` (2CQ) waits for that trace's completion event on the host and reads on CQ1, +so CQ0 can already run the next segment. + +**Segments.** Split a long graph into variants run back to back; hand intermediates from segment k to k+1 through an +in-place state (`ctx.write_state("mid", t)` in seg1, `ctx["mid"]` in seg2): a variant cannot see another variant's +outputs at warm-up time, a state buffer exists from the start. + +**Ping-pong state.** `pingpong=True` allocates two buffers and captures two traces per variant (phase 0 reads A / +writes B, phase 1 the reverse). A variant that writes the ping-pong states must write all of them ("stepping"); the +phase flips after each run of a stepping variant; read-only variants follow the current phase. In-place states +(default) read and write one buffer, so read what you need before the write. Reset with `reset_state(name, value)` +(host data, or a device tensor copied on the device). `write_state` traces one `ttnn.copy`; compute the new state +straight into `ctx.write_target(name)` (`ttnn.add(old, x, output_tensor=ctx.write_target("mem"))`, then +`ctx.write_state("mem", new)`) and no copy is traced (probe P13: one program less per state per frame). +`write_state` / `write_output` check shape, layout (ttnn.copy keeps it) and dtype changes (TILE only) before tracing. +For a ring buffer use `ttnn.copy` / slices into a fresh tensor and `write_state`; `ttnn.experimental.slice_write` +into a TILE state re-binds the handle and breaks traces (probe P13). + +**Per-stream state (PLAN.md D16).** A trace bakes one set of state addresses, so several `stream.id`s share them: +`add_state(..., banks=S_MAX)` allocates `S_MAX` DRAM banks of the state before capture, `save_state(name, i)` / +`load_state(name, i)` copy between the state and bank `i` (eager `ttnn.copy` on CQ0, ordered after the enqueued +replays; ~80 us for a StreamPETR-size state), and `StreamBanks` is the policy every temporal bundle uses: + +```python +from .ttaw.trace import StreamBanks +runner.add_state("prev_bev", shape=(1, 1, 22500, 256), dtype="bfloat16", pingpong=True, + banks=S_MAX if S_MAX > 1 else 0) # S_MAX = 1 at the first publish +streams = StreamBanks(runner, ["prev_bev"], max_streams=S_MAX, on_full="reject", max_gap_s=2.0) +fresh = streams.select(stream.get("id", "default"), reset=stream.get("reset", False), + timestamp_s=stream.get("timestamp_s")) # in _prepare, under the model lock +params = {"use_prev_bev": 0.0 if fresh else 1.0} # first-frame RT-dev params +``` + +`select` returns True when the stream starts fresh (new id, `reset=True`, timestamp backwards or a gap above +`max_gap_s`) after resetting its states to `init`; a new id beyond `max_streams` raises `InputError` (HTTP 400) with +`on_full="reject"` or takes the least recently used stream's bank with `"evict"` (`S_MAX=1`: restart the one state). +`forget(id)`, `streams` (most recent first), `active`, `describe()` (for `model.info`). + +**Outputs.** A variant returns either the tensors its last ops produce (allocated during the capture, alive as long +as the trace) or persistent outputs: `add_output(name, shape=...)` allocates a buffer before any capture and +`ctx.write_output(name, value)` copies into it inside the trace and returns it. Persistent outputs keep their address +across recaptures, can be shared by several variants (shape buckets with one readback) and are never overwritten by +another variant's replay; they cost one `ttnn.copy` per output. Op-produced outputs of a variant captured later can +be overwritten when an earlier-captured variant replays: read them before running another variant (`read` and +`__call__` return copies), and give segments read on CQ1 while CQ0 already replays the next one persistent outputs. +Warm-up and release free op outputs without `force`, so an output that shares memory with something the runner +does not own (a view of a weight) is never freed by the runner. + +| `TraceRunner` member | signature / meaning | +|---|---| +| constructor | `TraceRunner(device, *, num_command_queues=None, warmup_runs=1, stage_inputs=False, forbid_cache_misses=True, alloc_tracking=None, name="model")` | +| `add_input` | `(name, init=None, *, shape=None, dtype="bfloat16", layout=ROW_MAJOR, memory_config=DRAM, stage_fn=None)` -> device tensor. `init`: numpy / torch / ttnn host tensor (warm-up data), else zeros | +| `add_param` | `(name, value=0.0, *, shape=(1,1,1,1), dtype="float32", layout=TILE, memory_config=DRAM)` -> device tensor; broadcasts in ttnn binary ops | +| `add_state` | `(name, init=None, *, shape=None, dtype="float32", layout=TILE, memory_config=DRAM, pingpong=False, banks=0)` -> buffer or (A, B) | +| `add_output` | `(name, *, shape, dtype="float32", layout=TILE, memory_config=DRAM)` -> persistent output buffer (zeros) | +| `add_variant` | `(name, fn, *, warmup_runs=None)`; `fn(ctx)` returns a device tensor, a `Packed`, or a list / tuple / dict of them | +| `add_eager_warmup` | `(fn)`: eager device work the model runs between replays; `fn()` runs in the first `capture()`, before any trace (refused after it) | +| `capture()` | warm-up + capture of all pending variants (idempotent) | +| `run(variant=None, inputs=None, params=None)` | upload + non-blocking replay; returns the device outputs | +| `read(variant=None, *, cq_id=0, as_torch=False)` | blocking read into preallocated host tensors -> numpy (or torch); `Packed` -> `{name: array}` | +| `__call__(variant=None, inputs=None, params=None, *, as_torch=False)` | `run` + `read` | +| `upload(inputs=None, params=None)` / `replay(variant=None, n=1)` | the two halves of `run` (benches, pipelining) | +| `run_eager(variant=None, inputs=None, params=None)` | same function without a trace, buffers freed before returning: replay-vs-eager bit checks | +| `set_params(**values)` / `write_input(name, value)` | queue param values for the next run / upload an input now (CQ0; with 2 CQs the event the next CQ1 upload waits for is re-recorded after it) | +| `reset_state(name=None, value=None)` / `read_state(name)` / `state_buffer(name)` / `phase` | state control (on CQ0, ordered after enqueued replays); `value` may be a device tensor | +| `save_state(name, bank)` / `load_state(name, bank)` | state <-> bank copies on CQ0 (`add_state(banks=...)`; D16) | +| `outputs(variant=None, phase=None)`, `trace_ids()`, `describe()`, `timings_ms`, `phases`, `captured` | introspection (`describe()` is JSON-able: put it in `model.info`) | +| `release()` | release traces and persistent tensors (idempotent; also `__exit__`) | + +`TraceContext` (the `ctx` of a variant): `ctx[name]` (input, param, persistent output, or a state's read buffer), +`ctx.state(name)`, `ctx.write_target(name)`, `ctx.write_state(name, value)`, `ctx.write_output(name, value) -> +buffer`, `ctx.variant`, `ctx.phase`, `ctx.capturing`, `ctx.device`. Registering inputs / states / variants from +inside a variant function (during warm-up or capture) raises. + +`StreamBanks(runner, states, *, max_streams=1, on_full="reject", max_gap_s=None)`: `.select(stream_id, *, +reset=False, timestamp_s=None) -> fresh`, `.forget(id)`, `.streams`, `.active`, `.describe()` (above). + +`pack_outputs(tensors: dict, *, dtype="float32", align=32, row_elems=None) -> Packed`: one ROW_MAJOR tensor holding +every output (typecast to `dtype`; float32 is exact for bf16 and integers < 2**24), read with ONE device-to-host copy; +`TraceRunner.read` unpacks it, `Packed.layout.unpack(flat)` gives `{name: array}` (views) from any array of +`layout.total` elements (`PackLayout(entries, total, rows, row_elems)`, `.shape` = the device shape; `PackEntry(name, +offset, numel, shape, pitch=0)` indexes the elements in row-major order; `pitch > 0`: the tensor's rows of +`shape[-1]` elements are stored `pitch` elements apart, zero-padded, and `unpack` / `entry.view(flat)` return a +strided view). Layout (since 0.15.0, YOLOX PORT_LOG Q9): + +- packed total <= `SINGLE_ROW_MAX_ELEMS` (131072 = 512 KiB of fp32): one `[1, 1, 1, total]` row, each tensor + flattened and zero-padded to `align` elements -- exactly the layout and programs of 0.1.0-0.14.0; +- above it, or with `row_elems=R` (a multiple of 32): `[1, 1, rows, R]`, R = `PACK_ROW_ELEMS` (1024) by default. + Each tensor starts on a row boundary and is zero-padded to whole rows (a tensor of at most 64 KiB is flattened + and padded by < R elements; a larger `[.., n, c]` gets zero rows appended until `n * c` fills whole rows of R, + i.e. n rounded up to a multiple of `R / gcd(c, R)`; when that wastes more than 1/8 of the tensor beyond the best + alternative -- few rows of an awkward width: 0.15.0-0.18.0 packed `[1, 1, 2, 20001]` fp32 as 80 MB -- it is + flattened if its padded row fits 128 KiB, else its rows are zero-padded to a pitch `p` = c rounded up to a power + of two dividing R, or to R, the smallest segment: one `ttnn.pad` of the row and last dims, then the reshape into + rows of R; segments stay below 1.4x the tensor + one row, typical shapes keep the contiguous layout), then the + segments are concatenated on the row dim. No ROW_MAJOR page above 128 KiB is reshaped, so outputs of tens of MB + pack; a flat vector `[1, 1, 1, N]` wider than that is cut into `ttnn.slice` chunks; any other tensor with rows + wider than 128 KiB raises `ValueError` (give it a narrower last dim). Measured (p150b, + `tests/device/test_pack_outputs_device.py`, two runs, 1CQ / 2CQ; pack = replay of the packing trace, read = D2H + + unpack): 64 KB single row 0.15-0.22 + 0.14-0.18 ms; 1 MB (YOLOX 14400 x 13 fp32 head level + a 240 x 240 uint32 + map) 0.69-0.75 + 0.30-0.54 ms; 16 MB (`[1, 1, 262144, 16]` fp32) 1.24-1.74 + 3.7-8.1 ms (2-4 GB/s); bit-exact + through trace replays (`logs/ttaw/q9_pack_outputs_device*.log`). + Why: a single ROW_MAJOR row is ONE page, and `ttnn.reshape` / `concat` stage whole pages in L1 (2 x the + destination page per kernel copy for pages that are not 16-byte aligned), so the 0.14.0 layout failed above + ~0.6 MB with `TT_FATAL: RM reshape dest staging does not fit in L1`. + +`CQ_COMPUTE = 0`, `CQ_INPUT = 1`. + +**Alloc tracking (debug).** `TT_METAL_TRACE_ALLOC_TRACKING=1` must be exported before Python imports ttnn; then +`ttnn.execute_trace` refuses to replay while a buffer allocated after a capture is alive. `alloc_tracking_enabled()` +reports it; `TraceRunner(alloc_tracking=True)` raises if it is off. The runner acknowledges the outputs of every trace +captured after the first (plain ttnn would flag them: `tests/device/test_trace_alloc_tracking_device.py`). + +--------------------------------------------------------------------------------------------------------------------- + +## 4. Tensors (C03, `ttaw.tensors`) + +| name | meaning | +|---|---| +| `TILE`, `round_up(n, multiple=32)`, `tile_padded_shape(shape)`, `pad_to_tile(a, value=0)` | tile geometry (host) | +| `pad_to_capacity(a, capacity, *, axis=0, value=0, truncate=False) -> (padded, n_valid)` | fixed-capacity buffers (capacities are grid-independent constants) | +| `as_4d(a)` | prepend unit dims to rank 4 | +| `float32_to_bf16_bits`, `bf16_bits_to_float32`, `round_to_bf16` | RNE bf16 conversion identical to torch (expected values for exact tests) | +| `ttnn_dtype(name)`, `dtype_name(dtype)` | `"bf16"`, `"fp32"`, `"bfp8"`, `"bfp4"`, `"uint32"`, `"int32"`, `"uint16"`, `"uint8"` aliases | +| `to_host_tensor(value, dtype, layout=None, *, shape=None)` | numpy / torch / ttnn host tensor -> ttnn host tensor; integers never pass through a float intermediate (int64 input to `ttnn.from_torch` would) | +| `to_device(value, device, dtype="bf16", layout=None (TILE), memory_config=None (DRAM))` | eager upload | +| `to_numpy(t)` | ttnn (device or host) / torch -> numpy (bf16 -> float32) | +| `to_layout(t, layout)`, `typecast(t, dtype)`, `to_fp32(t)`, `to_bf16(t)` | no-ops when nothing changes (no extra program) | +| `fp32_island(fn, *tensors, out_dtype="bfloat16")` | run `fn` on fp32 copies, cast results back | +| `upload_u8(array, device, *, out_dtype="bfloat16", layout=None, memory_config=None)` | uint8 upload + on-device typecast (exact); in a trace keep the uint8 tensor as the input and `ttnn.typecast` as the first op | +| `HostStaging(shape, dtype)` | persistent ROW_MAJOR host tensor; `.write(array)` copies into its buffer through a `torch.from_dlpack` alias (`zero_copy=True` on this tree for float32 / bfloat16 / uint32 / int32 / uint16 / uint8) and returns `.tensor`; pass it as a `run(inputs=...)` value | + +--------------------------------------------------------------------------------------------------------------------- + +## 5. Weights (C04, `ttaw.weights`) + +```python +from .ttaw.weights import OnnxWeights, WeightCache, fold_bn_conv + +w = OnnxWeights(weights_path / "pts_backbone_neck_head_centerpoint.onnx") # parsed as data, never executed +for node in w.nodes("Conv", "/backbone/*"): # graph order, glob or regex + wf, bf = w.fold_conv_bn(node.name) # + its BatchNormalization consumer, fp64, rounded once +k = w.param("/dit/blocks.0/attn/MatMul", 1) # anonymous initializer, addressed by consuming node + slot +cache = WeightCache("tt_centerpoint/base", w.sha256, version=f"{__version__}.prep1") # $TT_CACHE_PATH | ~/.cache/ttaw +tw = cache.get("backbone.0.conv.w", lambda: wf, dtype="bfloat16", layout=ttnn.ROW_MAJOR_LAYOUT, device=dev) +``` + +| name | meaning | +|---|---| +| `OnnxWeights(path, *, load_external_data=True)` | `.array(name)` (initializer / Constant, through Identity), `.has`, `.initializer_names()`, `.state_dict()`, `.node(name)`, `.nodes(op_type=None, pattern=None, *, regex=False)`, `.find_node(pattern, op_type=None, *, regex=False)`, `.producer(t)`, `.consumers(t)`, `.consumer_of(t, op_type=None)`, `.param(node, slot)`, `.params(node)`, `.conv(node) -> ConvParams`, `.gemm(node) -> GemmParams`, `.matmul_weight(node)`, `.batchnorm(node) -> BatchNorm`, `.fold_conv_bn(conv, bn=None, *, dtype=np.float32)`, `.sha256`, `.input_names`, `.output_names`, `.opset`. External data outside the model directory is refused. Unnamed nodes are `"_"` | +| `OnnxNode`, `ConvParams` (weight, bias, strides, pads, dilations, group, kernel_shape, output_padding, auto_pad, output_shape), `GemmParams` (weight as stored, bias, trans_a, trans_b, alpha, beta) | frozen records; `pads` is the explicit attribute: with `auto_pad` `SAME_*` derive the padding from the input size | +| `BatchNorm(gamma, beta, mean, var, eps=1e-5)` | float64; `.scale`, `.shift`, `.channels`, `.apply(x, axis=1)`, `.from_state_dict(sd, prefix, eps)` | +| `fold_bn(w, b, bn, *, axis=0, dtype=np.float32)` | generic fold (`dtype=None` keeps float64 for a single final rounding to the device dtype) | +| `fold_bn_conv(w, b, bn)` | Conv weight `[Cout, Cin/g, k...]`, any groups | +| `fold_bn_linear(w, b, bn, *, layout="out_in")` | torch Linear / Gemm transB=1 (`out_in`) or MatMul / Gemm transB=0 (`in_out`); also BN1d | +| `fold_bn_conv_transpose(w, b, bn, *, groups=1)` | ConvTranspose weight `[Cin, Cout/g, k...]` | +| `load_safetensors(path)`, `load_torch_checkpoint(path, *, key=None)` (`torch.load(weights_only=True)`), `load_state_dict(path)` -> `StateDict` | `.safetensors`, `.pth/.pt/.ckpt/.bin`, `.npz` | +| `StateDict(tensors)` | `.sub(prefix)`, `.strip(prefix)`, `.bn(name, eps)` | +| `WeightCache(namespace, source_digest="", *, version="", root=None, enabled=None)` | `.get(key, make, *, dtype, layout=None, device=None, memory_config=None, shape=None)` caches ttnn host tensors as `.tensorbin` (atomic writes; rebuilt if unreadable or of another dtype / layout / `shape`), `.path(...)`, `.clear()`, `.hits`, `.misses`; `TTAW_WEIGHT_CACHE=0` disables. The cache outlives images (`/tensor-cache` is the host's `~/.cache/tt-model//tensors`): pass `version` (bundle `__version__` + a prep revision) and bump it whenever the code that prepares the tensors changes | +| `cache_root()`, `file_sha256(path)` | helpers | + +--------------------------------------------------------------------------------------------------------------------- + +## 6. Precision and knobs (C05, `ttaw.precision`, `ttaw.knobs`) + +`ttnn.matmul` / `linear` drop to LoFi when a `program_config` or `core_grid` comes without a compute config, and +`ttnn.WormholeComputeKernelConfig()` without `math_fidelity` is `MathFidelity.Invalid`. Always pass one: + +```python +from .ttaw.precision import PrecisionPolicy, compute_kernel_config +POLICY = PrecisionPolicy({"head.reg*": "accurate", "backbone.*": "HiFi2+fp32:w=bfp8"}, default="balanced") +policy = POLICY.with_env("CENTERPOINT") # CENTERPOINT_PRECISION="backbone.*=HiFi4+fp32;*=HiFi2+fp32" +y = ttnn.linear(x, w, compute_kernel_config=policy.compute_kernel_config("head.reg.fc1"), program_config=pc) +w_dtype = policy.resolve("backbone.block3.conv2").weights_dtype() +``` + +| name | meaning | +|---|---| +| `compute_kernel_config(fidelity="HiFi2", *, fp32_acc=True, approx=False, packer_l1_acc=False, dst_full_sync=False)` | explicit `ttnn.WormholeComputeKernelConfig` (= `BlackholeComputeKernelConfig`) | +| `Precision(fidelity="HiFi2", fp32_acc=True, approx=False, packer_l1_acc=False, dst_full_sync=False, weights="bfloat16", activations="bfloat16")` | `.parse("HiFi4+fp32+approx+l1acc+fullsync:w=bfp8:a=bf16" or preset)`, `.label`, `.with_(**)`, `.compute_kernel_config()`, `.weights_dtype()`, `.activations_dtype()` | +| `PRESETS` | `accurate` (HiFi4 + fp32), `balanced` (HiFi2 + fp32, the default), `fast` (LoFi, bfp8 weights; only after gates pass) | +| `PrecisionPolicy(rules, default="balanced")` | first matching glob wins; `.resolve(module)`, `.compute_kernel_config(module)` (cached), `.override(spec)`, `.with_env(prefix, env=None)`, `.describe()` (rules + which modules resolved to what) | +| `FIDELITIES` | `("LoFi", "HiFi2", "HiFi3", "HiFi4")` | + +Other silent defaults to override by hand: `ttnn.layer_norm` epsilon 1e-12, SDPA `is_causal=True`, `ttnn.embedding` +PADDED returning the cached pad row, HARDSWISH fused into conv2d skipped. + +Knobs (one per optimization, default = measured best, pinned in `tt-model.yaml serve.env`): + +```python +from .ttaw.knobs import Knob, Knobs +KNOBS = Knobs("CENTERPOINT", [Knob("FUSED_HEAD", True, "merged head convs"), + Knob("ACT_BLOCK_H", 64, "conv act_block_h", choices=(32, 64, 128)), + Knob("BFP8_WEIGHTS", False, "bfp8 backbone weights", experiment=True)]) +knobs = KNOBS.read() # once, in _build; env CENTERPOINT_FUSED_HEAD=0 is the A/B switch +assert KNOBS.serve_env() == yaml_serve_env_subset # host test: the image pins the defaults +``` + +`Knob(name, default, doc="", choices=None, experiment=False)`; `Knobs(prefix, knobs)`: `.read(env=None, +**overrides) -> KnobValues` (attribute / item access, `.source(name)`, `.overridden()`, `.as_dict()`, immutable; +warns when an experiment knob is set), `.defaults()`, `.serve_env(values=None)`, `.doc_table()`, `.env_name(name)`; +`parse_bool(text)`. + +--------------------------------------------------------------------------------------------------------------------- + +## 7. Metrics, goldens, gates (C06, `ttaw.metrics`, `ttaw.golden`) + +```python +# tests/test_pcc_device.py of a bundle +from ..ttaw.golden import GateRegistry, compare_taps, load_goldens +from ..ttaw.metrics import pcc, match_detections + +GATES = GateRegistry.for_test(__file__, {"backbone": 0.999, "head.heatmap": 0.99, + "dets.recall": (0.95, "min", "recall"), "plan.ade": (0.5, "max", "ade")}) + +def test_taps(device_taps): # numpy dict from an eager device run or trace outputs + with load_goldens(SPEC_DIR / "golden/sample0.npz") as gold: + report = compare_taps(device_taps, gold, gates=GATES, names=["backbone", "head.heatmap"]) + report.save_json(LOGS / "compare_backbone.json"); print(report.table()) + assert report.passed +``` + +`GateRegistry` stores frozen gates in `.gates.json` next to the test at the first green run. A later +declaration that is looser (lower `min`, higher `max`, other direction or metric) raises `GateLoosenedError`; +tighter ones are re-frozen; failing checks freeze nothing. `TTAW_GATES_READONLY=1` (or `write=False`) never writes. +Changing a frozen gate = editing the JSON by hand + disclosure in VERIFICATION_.md. + +| name | meaning | +|---|---| +| `pcc(t, r)` | float64 Pearson; NaN on non-finite input; a constant side (detected exactly, max == min) gives 1.0 for two equal constants (`np.allclose`) and 0.0 otherwise, also for two different constants | +| `masked_pcc(t, r, mask)`, `valid_row_pcc(t, r, n_valid=None, *, rows=None, axis=0)` | padded / masked buffers | +| `error_stats(t, r)` | `{pcc, max_abs, mean_abs, rel_l2, n}` | +| `argmax_agreement(t, r, *, axis=-1, mask=None)`, `label_agreement(t, r, *, mask=None)` | class decisions (ties -> lowest index) | +| `mask_iou(a, b)`, `mean_iou(t, r, num_classes, *, ignore_index=None) -> (miou, per_class)`, `box_iou_xyxy(a, b)` | IoU | +| `topk_overlap(t_idx, r_idx, k=None)`, `topk_set_overlap(t_scores, r_scores, k)` | data-dependent selections | +| `match_detections(t_centers, t_labels, t_scores, r_centers, r_labels, r_scores, *, max_dist=0.5, dims=2, same_label=True) -> DetectionMatch` | greedy same-label BEV-centre matching (`.pairs`, `.recall`, `.precision`, `.matched`, `.center_errors`, `.score_errors`, `.to_dict()`); box arrays `[N, 7]` work as centres; a box with a non-finite centre never matches | +| `ade_fde(pred, ref, *, dims=2) -> (ade, fde)` | trajectories `[..., T, D]` | +| `as_array(x)` | numpy / torch (bf16 ok) / list -> numpy | +| `save_goldens(path, tensors, meta=None, *, compress=False)`, `load_goldens(path) -> Goldens` | `.npz` + JSON `__meta__`; `Goldens` is a lazy mapping with `.meta`, `.sha256`, `.close()`, context manager | +| `TapRegistry(enabled=True, *, include=("*",), exclude=())` | `.tap(name, value) -> value` (device tensors read at once; raises inside a capture), `.scope(prefix)`, `.wants`, `.names()`, `.to_dict()`, `.save(path, meta)`, `.clear()`; `NULL_TAPS` is a disabled registry | +| `Gate(threshold, direction="min", metric="pcc")` | `.of(spec, metric=None)` (a bare number takes the direction of `metric` from `DIRECTIONS` and is refused for a metric of unknown direction; a direction contradicting a known metric, e.g. a `"max"` PCC gate, is refused), `.passes(v)` (NaN fails), `.looser_than(other)` | +| `GateRegistry(path, gates=None, *, write=None)` | `.for_test(test_file, gates)`, `.check(name, value, gate=None) -> GateResult`, `.require(...)` (AssertionError), `.gate(name)`, `.names()`, `.frozen`, `.results`, `.save()` | +| `CompareReport(title="", meta=None, gates=None)` | `.add(name, test, ref, *, metric="pcc", gate=None)`, `.add_value(name, value, *, metric, gate=None, stats=None)`, `.passed`, `.table()`, `.to_dict()`, `.save_json(path)`; `TapComparison` records | +| `compare_taps(test, golden, *, gates=None, metric="pcc", names=None, title="", meta=None)` | one report over common (or named) taps; missing taps fail when gated | +| `METRICS` | name -> (function, direction) used by reports (`pcc`, `masked_pcc`, `argmax_agreement`, `label_agreement`, `mask_iou`, `max_abs`, `mean_abs`, `rel_l2`) | +| `DIRECTIONS` | metric name -> `"min"` / `"max"` for every gate metric that may be given as a bare number: the `METRICS` plus `recall`, `precision`, `iou`, `miou`, `agreement`, `topk_overlap` (min) and `ade`, `fde`, `center_err*`, `score_err*`, `max_abs_err` (max) | + +--------------------------------------------------------------------------------------------------------------------- + +## 8. Bench and profile (C07, `ttaw.profiling`) + +```python +from .ttaw.profiling import AiclkSampler, bench_trace_runner, signposted, summarize_ops + +with AiclkSampler(interval_s=0.05) as clk: # sysfs tt_aiclk + hwmon power / temperature + bench = bench_trace_runner(model.runner, "default", {"pillars": p}, iters=100, post=model.decode) +bench.save_json(f"{ROOT}/logs/centerpoint/bench_baseline.json", config=model.device_info, aiclk=clk.summary()) +print(bench.table()) # host_in / h2d / trace / d2h / post / e2e / b2b, p50 p99 + +with signposted("trace"): # profile_ops.py under `python -m tracy -r -p -v -o ...` + model.runner.replay("default"); ttnn.synchronize_device(dev) +print(summarize_ops(f"{ROOT}/generated/profiler/centerpoint_baseline").table()) # ops, kernel sum, op2op, span +``` + +| name | meaning | +|---|---| +| `StageBench(name="")` | `.stage(name)` context, `.add(name, ms)`, `.summary()` (`{stage: {n, p50, p99, mean, min, max}}`), `.table()`, `.to_dict(**extra)`, `.save_json(path, **extra)`; `STAGES` order | +| `bench_trace_runner(runner, variant, inputs, *, params=None, iters=100, warmup=10, post=None, b2b_iters=None, name="")` | the standard stage breakdown of a `TraceRunner` variant | +| `time_b2b(enqueue, sync, n=100, warmup=5) -> ms` | back-to-back device time per iteration | +| `signpost(name, message=None) -> bool`, `signposted(name)` | Tracy markers (no-op without `tracy`) | +| `read_device_profiler(device)` | `ttnn.ReadDeviceProfiler` per trace segment (the buffer holds ~1000 ops) | +| `summarize_ops(source, *, start="trace", end=None, last_replay_session=None, freq_mhz=None, top_gaps=10) -> OpsSummary` | `ops_perf_results*.csv` / `cpp_device_perf_report.csv` / directory / rows; `OpsSummary`: `ops`, `kernel_sum_us`, `fw_sum_us`, `op2op_sum_us`, `span_us`, `by_op`, `fidelity`, `top_gaps`, `.table(top)`, `.to_dict()`. CLI: `python -m .ttaw.profiling [--start trace] [--end trace_end] [--json out]` | +| `find_ops_csv(path)` | newest ops CSV under a directory | +| `AiclkSampler(chip=0, interval_s=0.05, *, root="/sys/class/tenstorrent")` | `.start()`, `.stop()`, context manager, `.sample_once()`, `.summary()`, `.available` | + +`common/tools/p2_trace_bench.py --mode eth-1cq|eth-2cq|eth-2cq-staged|worker-2cq` measures the runner itself. + +--------------------------------------------------------------------------------------------------------------------- + +## 9. I/O and outputs (C08, `ttaw.io`, `ttaw.outputs`) + +`ttaw.io` (numpy only; every client mistake raises `InputError`, which the server maps to HTTP 400): +`b64decode(s, *, field, max_bytes)`, `PointCloud(points, fields, frame_id)` (`.select(names, fill=None)`; rows with a +non-finite `x` / `y` / `z` column -- the first three columns when unnamed -- dropped), +`decode_points(spec, *, max_points=2_000_000, max_bytes=None, default_fields=DEFAULT_POINT_FIELDS)`, +`load_points(source, *, fields=None, fmt=None, frame_id="base_link", default_fields=...)` (path `.bin/.npy/.npz/.pcd`, +bytes + `fmt`, arrays, structured arrays, envelopes), `decode_image(raw, *, fmt="auto", max_side=8192)`, +`load_image(image)` -> RGB uint8 HxWx3, `CameraImage`, `decode_cameras(images, calibration=None, *, order=None, +require_calibration=True, max_bytes=None)` (the server's `images[]`), its Python-API twins (0.15.0) +`load_camera(source, *, name=None, calibration=None, require_calibration=False)` (a `CameraImage`, `{"camera", +"image" | "path" | "data", "intrinsics", "T_ref_from_camera", "distortion", "timestamp_s"}` or any `load_image` source +with `name`; inline calibration wins over `calibration["cameras"][name]`) and `load_cameras(images, calibration=None, +*, order=None, require_calibration=True, calib_dir=None)` (a list or a `{name: image}` mapping; `{"preset": name}` +resolved in `calib_dir`; reordered to `order`, every camera exactly once: `model(images=..., calibration=...)` and +`/predict` see the same cameras), `decode_rois(rois, *, cameras=None, labels=None, max_rois=4096)` (2-D detections +as an input, `[{"camera", "label", "score", "box_xyxy"}]`: names / indices checked against `labels`, score in [0, 1], +finite ordered boxes; PointPainting `rois`, BUNDLE_CONVENTIONS.md 7.2), `parse_transform(obj, *, field)` (4x4 / 3x4 / `{matrix}` / +`{translation, rotation_wxyz|rotation_xyzw|rotation}` / `{x, y, z, roll, pitch, yaw}` tf2 RPY), +`parse_intrinsics(obj, *, field)`, `resolve_calibration(calibration, calib_dir)` (`{"preset": name}` -> +`calib_dir/.json` merged), `decode_named_arrays(spec, schema=None, *, max_bytes=None)` (planner inputs), +`check_named_arrays(arrays, schema)` (names, shapes, numeric and finite values, integral values in range for integer +dtypes; every problem is an `InputError`), `load_named_arrays(source, schema=None)` (mapping / envelope / `.npz` +path or bytes: the Python-API form), +`encode_array(arr, *, fmt="npz"|"npz_compressed"|"raw"|"list", key)`, `encode_png(image, *, key)`, `to_jsonable(obj)` +(float32 / float16 scalars rounded to 6 significant digits, float64 exact), +`POINT_FORMATS`, `DEFAULT_POINT_FIELDS = ("x", "y", "z", "intensity")`, `MAX_POINTS_DEFAULT`. + +`ttaw.outputs` (each has `to_dicts()` and `to_dict(output_format="json"|"npz")` = the `/predict` body with `model`, +`frame_id`, `meta`, `timing_ms`): + +| class | fields | +|---|---| +| `Detections3D(boxes [N,7] x y z l w h yaw, scores, label_ids, velocities=None, labels=(), model="", frame_id="base_link", timing_ms, meta)` | rows sorted by score at construction; `detections[] {label, label_id, score, center, size, yaw, velocity}`; npz adds `arrays` | +| `Detections2D(boxes_xyxy, scores, label_ids, labels=(), extras={}, model, frame_id="camera", ...)` | `detections[] {label, label_id, score, box_xyxy}`; `extras` (e.g. `{"semseg": Mask2D}`) encoded by name | +| `Segmentation3D(label_ids [N], class_names=(), scores=None, ...)` | `labels` (npz uint8 / uint16), `class_names`, `class_counts`, `scores` | +| `Mask2D(mask (H, W), class_names=(), encoding="png", ...)` | `mask` (png, or npz for `output_format="npz"` / non-uint8), `class_counts`; `.payload(fmt)` | +| `Trajectory(poses [T, D], columns=("x", "y", "yaw"), turn_indicator=None, predicted_agents=None, ...)` | `trajectory`, `columns`, `num_poses`, `turn_indicator`, `predicted_agents` (npz) | + +`label_name(labels, id)` maps ids to names (the id as text when out of range). + +--------------------------------------------------------------------------------------------------------------------- + +## 10. `ModelBase` (C08, `ttaw.api_base`) + +```python +from .ttaw.api_base import ModelBase +from .ttaw.outputs import Detections3D + +class CenterPoint(ModelBase): + MODEL_NAME = "centerpoint-p150"; ENV_PREFIX = "CENTERPOINT" + DEFAULT_REPO = "AutowareFoundation/lidar_centerpoint"; DEFAULT_TAG = "v4.1" + DEFAULT_REVISION = "494c8171def40bd36cc2feb323e0a5acbfab132b"; ALLOW_PATTERNS = ["base/*", "tiny/*"] + VARIANTS = ("base", "tiny"); DEFAULT_VARIANT = "base"; INPUT_KIND = "lidar" + LABELS = ("CAR", "TRUCK", "BUS", "BICYCLE", "PEDESTRIAN") + RUNTIME_PARAMS = {"score_threshold": (float, 0.0, 1.0, 0.35), "max_detections": (int, 1, 1000, 500)} + DEVICE_DEFAULTS = {"num_command_queues": 1, "trace_region_size": 64 << 20} + + def _build(self): # weights -> device tensors, TraceRunner + inputs / params / states / variants + self.runner = TraceRunner(self.device, name=self.MODEL_NAME); ... + def _warm_one(self, v): # capture (default_warmup_variants() -> [{"variant": self.variant}]) + self.runner.capture() + def _prepare(self, points=None, **kw): # host pre-processing; raise io.InputError for client mistakes + ... + def _forward(self, prep): # under the model lock + return self.runner("default", inputs={...}) + def _postprocess(self, raw, prep, params) -> Detections3D: + return Detections3D(boxes, scores, ids, labels=self.LABELS, model=self.MODEL_NAME) + def _release(self): + self.runner.release() + def extra_info(self): + return {"trace": self.runner.describe(), "knobs": self.knobs.as_dict()} + +with CenterPoint.from_pretrained() as model: # weights (pinned) -> device (ETH, 12x10) -> build -> warm + out = model("samples/test.npz", score_threshold=0.4) + print(out.to_dict()["num_detections"], out.timing_ms, model.info) +``` + +`from_pretrained(model_id=None, *, revision=None, variant=None, device_id=None, device=None, dispatch=None, +num_command_queues=None, weights_dir=None, warmup_variants="default", verbose=False, **compile_params)`: weights are +resolved first (`weights_dir` > `_WEIGHTS_DIR` > local dir `model_id` > HF snapshot at the pinned revision, +offline fallback to the cache), then the device opens with `DEVICE_DEFAULTS` < `_*` / `TT_DEVICE_ID` < explicit +arguments; a caller-provided `device` is left open by `close()`. Variant default: `_VARIANT` or +`DEFAULT_VARIANT`. A failing build closes everything. + +Other members: `resolve_weights(...)` (also module-level `resolve_weights(model_id, *, revision, allow_patterns, +weights_dir, env_var)`), `device_config(**overrides)`, `requires_calibration()`, `validate_params(params)` (unknown / +outside `[min, max]` (either bound may be None) / wrong type -> `InputError`; bools must be real booleans, ints must +be integral and not booleans, floats finite; numeric strings are accepted), `warmup(variants="default")` (idempotent), +`default_warmup_variants()`, `__call__(points=None, *, sweeps, images, calibration, stream, inputs, **params)` / +`predict` (one re-entrant lock around prepare + forward + postprocess; fills `timing_ms` preprocess / device / +postprocess / total), `info`, `closed`, `close()` (idempotent, also at interpreter exit), context manager. +Class attributes also: `WEIGHTS_LICENSE`, `CAMERA_ORDER`, `REQUIRE_CALIBRATION` (None: true for multicam), +`POINT_FIELDS`, `EXTRA_INPUTS` (model-specific input keywords handed to `_prepare` instead of being validated as +runtime params, e.g. PointPainting `("rois",)`; the server's `decode_extra` adds them to the call kwargs) and +`INPUT_SCHEMA` (`{name: (shape with None for free dims, dtype)}`: `inputs=` is decoded with +`io.load_named_arrays` and checked on every call, API and server alike). `RUNTIME_PARAMS` may not reuse an input +keyword (`from_pretrained` raises `TypeError`). `info` adds `extra_inputs` and `input_schema`; `__call__` re-checks +`closed` once it holds the lock. + +--------------------------------------------------------------------------------------------------------------------- + +## 11. Server (C08, `ttaw.server`) + +`create_app(spec: ServerSpec, *, model_factory=None) -> FastAPI` implements BUNDLE_CONVENTIONS.md section 7: +`GET /` (routes), `GET /health` and `/v1/health` (always 200: `ok` / `starting` / `error`, plus device), +`GET /info` (model, task, io, autoware, weights, device, input limits, calibration presets, labels, variants, +runtime / compile params, warm-up and boot ms, source), `GET /v1/models` (stub), `POST /predict`. Errors: 400 +(`InputError`, bad params, body above `_MAX_BODY_MB` by Content-Length), 422 (schema: unknown fields are +refused), 503 (starting / failed boot), 500 (`inference failed: : `). The lifespan reads the environment +once (`config_from_env`), refuses a `TT_MESH_SHAPE` other than 1x1, builds the model through +`spec.model_cls.from_pretrained` (or `model_factory(cfg)`), and closes it on shutdown under the lock. A body holding a +NaN / infinity (strict JSON has none) is a 500 `inference failed: non-finite value at body.`. `/info` `input` +also lists `extra_inputs` and `input_schema`. + +`ServerSpec(model_name, env_prefix, model_cls=None, task="", default_weights="", owner="changh95", hardware=HARDWARE, +io="", autoware={}, source={}, calib_dir=None, default_variant=None, version="0.1.0", description="", +request_model=PredictRequest, decode_extra=None, info_extra=None)`. Model-specific request fields: subclass +`PredictRequest` (e.g. `rois: Optional[List[...]] = None`), add them to the call kwargs in +`decode_extra(request, kwargs, max_bytes)` and list them in the model's `EXTRA_INPUTS`. `app.state.ttaw` +(`ServerState`) holds `model`, `config`, `ready`, `error`, `lock`, `model_factory` (tests: assign a stub before +`TestClient(app)`) and `predict` (the route handler, for API == server checks). Also exported: +`parse_mesh_shape`, `config_from_env(spec, env=None)`, `default_model_factory(spec)`, `decode_request(req, spec, +max_bytes)`, the request models `PointsSpec`, `SweepSpec`, `ImageSpec`, `StreamSpec`, `ArraysSpec`, +`PredictRequest`. + +`ttaw.server.client` and `ttaw.server.smoke` are standard-library only and import nothing from ttaw, so a bundle's +`smoke_test.py` (run by `container_smoke.sh` with the host's `python3`, which has no numpy) loads them by path: + +```python +HERE = Path(__file__).resolve().parent # code//server +sys.path.insert(0, str(HERE.parent / "ttaw" / "server")) +import client as ttaw_client +import smoke as ttaw_smoke +health = ttaw_client.wait_ready(url, wait_s=600) +info, fails = ttaw_client.check_service(url, expect_dispatch="eth", expect_grid="12x10") # PLAN 0.3 item 6 +pinned = ttaw_smoke.pinned_config(f"{staged}/tt_kernel_manifest.json", profile) # serve.env + profile +fails += ttaw_smoke.check_pinned(info, pinned, "CENTERPOINT") # /info runs the pins +code, body = ttaw_client.post(url, ttaw_client.build_request(points=str(SAMPLE))) +if ref := ttaw_smoke.find_reference(SAMPLE, pinned["profile"], info.get("variant")): # stored CPU output + metrics, more = ttaw_smoke.compare_with_reference(body, ttaw_smoke.load_json(ref), gates={"min_recall": 0.95}) + fails += more +fails += [f] if (f := ttaw_client.check_bad_request(url)) else [] +print(ttaw_client.smoke_line("centerpoint-p150", f"n={body.get('num_detections')}", fails)) +``` + +Functions: `build_request(points=None, fields=None, images=(), calib=None, calib_preset=None, inputs=None, +params=None, stream_id=None, reset=False, output_format="json")`, `post(url, payload, timeout=120)`, +`get_json(url, timeout=10)`, `wait_ready(base, wait_s=0, poll_s=5)`, `check_service(base, *, expect_dispatch="eth", +expect_grid="12x10") -> (info, failures)`, `check_bad_request(base, payload=None)`, `smoke_line(model, summary, +failures)`, `b64file`, `main` (CLI: `--points/--image/--calib/--calib-preset/--inputs/--param/--out/--url`). + +`ttaw.server.smoke` functions: + +| name | meaning | +|---|---| +| `compare_with_reference(body, reference, gates=None) -> (metrics, failures)` | served `/predict` body vs the stored CPU-reference body of the same input, by the reference's keys: `detections` (greedy same-label matching in descending reference score; 3-D rows by BEV centre distance, 2-D rows by IoU; recall, precision, max score difference), `trajectory` (ADE / FDE over x, y), every encoded array (integer: fraction of equal elements; float: max abs difference, equal infinities agree, a NaN fails), `model` / `frame_id` equal. Undecodable input is a failure, never an exception | +| `DEFAULT_GATES` | `max_center_dist` 0.5 m, `min_iou` 0.5, `min_recall` 0.95, `min_precision` 0.95, `max_score_err` 0.05, `min_label_agreement` 0.99, `max_abs_err` None (report only), `max_ade` 0.5, `max_fde` 1.0; `gates=` overrides by name (unknown names raise) | +| `find_reference(sample, *keys)` | first existing `..reference.json` (keys: serve profile, model variant), else `.reference.json`, else None | +| `pinned_config(manifest, profile=None)` | `{"profile", "env", "weights"}` of a staged `tt_kernel_manifest.json`: `serve.env` with the profile's env on top (tt-model's merge), default profile when None | +| `check_pinned(info, pinned, prefix)` | `/info` agrees with the pinned `_DISPATCH`, `_NUM_CQS`, `_VARIANT` and the weights revision (absent pins are not checked) | +| `decode_array(spec) -> Array(shape, dtype, values)` | numpy-free decoding of `io.encode_array` (npz, npz_compressed, raw, list; C or Fortran order, either byte order) and `io.encode_png` (8-bit grey / RGB(A), every PNG filter) | +| `describe_metrics(metrics)`, `load_json(path)`, `ARRAY_FORMATS` | helpers | + +The reference body is the `to_dict()` of the fp32 CPU reference on the shipped sample, stored next to it as +`.reference.json` (per profile or variant when the output depends on it). Synthetic samples may give no +detections, so agreement with that reference, not a detection count, is their smoke gate (PLAN.md section 6.3). + +--------------------------------------------------------------------------------------------------------------------- + +## 12. Vendoring (C09, `common/tools/vendor.py`) + +`python common/tools/vendor.py [--pkg tt_] [--rev REV] [--allow-dirty] [--check] [--dry-run]`. + +- **What is vendored is a committed revision**: `ttaw/` as committed at `--rev` (default `HEAD`), read with `git + archive`, never the working tree of `common/` (other agents' uncommitted modules live there: YOLOX PORT_LOG Q7). + Uncommitted changes under `common/ttaw` (staged, unstaged, untracked) are listed in a WARNING and left out: + commit your change (`git add `, bump the version + CHANGELOG), then vendor. The revision is resolved + once, so a commit landing meanwhile never mixes into the recorded `source_commit`. +- `--allow-dirty` vendors the working tree instead, for a local experiment only: `source_dirty` is true when it had + uncommitted changes or its files differ from `ttaw/` of HEAD (.gitignore'd files included: `source_dirty: false` + always means "verified against `source_commit`"); `check_bundle.py` warns about such a copy and refuses it at + `--stage publish`. Outside a git checkout the working tree is vendored with `source_commit: null` (same + treatment). +- The destination must be a plain directory: one that is, or reaches through a symlink, `common/ttaw` (or an + ancestor of it) is refused before anything is written (mirroring HEAD there would delete other agents' untracked + files and revert their uncommitted changes). +- The copy mirrors that tree into `code//ttaw` (stale files removed; `__pycache__`, `*.pyc`, `*.bin`, `*.log` + never copied) and writes `VENDORED.json`: `version` (of the vendored `__init__.py`), `source_commit` (full sha), + `source_ref` (the `--rev` given), `source_tree` (`commit` | `working-tree`), `source_dirty`, `vendored_at`, + per-file sha256. `research/packaging/scripts/instantiate_bundle.py --out` vendors the same way (`--vendor-rev`, + `--vendor-allow-dirty`; an existing copy is kept unless `--force`, and always with `--only `). +- `--check` (exit 1 on any line): "modified in bundle" / "missing in bundle" / "extra in bundle" (the copy vs its + `VENDORED.json`); "not at recorded revision " (`VENDORED.json` is not `ttaw/` of `source_commit`; + "unverified:" when that commit is not in `common/`; skipped for a `source_dirty` copy); "outdated:" (`VENDORED.json` + vs `ttaw/` committed at `--rev`, default HEAD: re-vendor and re-run the gates). Uncommitted work in `common/ttaw` + never makes a bundle look outdated. `check_bundle.py` reports "outdated:" and "unverified:" as warnings (errors at + `--stage publish`), every other line as an error. + +Never edit the vendored copy: change `common/ttaw`, bump `__version__` + CHANGELOG, commit, re-vendor every consumer +and re-run their gates (PLAN.md section 5.1 step 5). Functions: `vendor(bundle, pkg=None, *, dry_run=False, rev=None, +allow_dirty=False)`, `check(bundle, pkg=None, *, rev=None)`, `load_source(rev=None, *, allow_dirty=False) -> Source`, +`rev_files(rev)`, `rev_hashes(rev)`, `rev_version(rev)`, `resolve_rev(rev)`, `dirty_paths()`, `tree_hashes(root)`, +`source_version()` (working tree), `git_state()`, `find_package(bundle, pkg)`; git calls never take the index lock +(`GIT_OPTIONAL_LOCKS=0`). + +--------------------------------------------------------------------------------------------------------------------- + +## 13. Measured facts and pitfalls (this p150b, tt-metal 44d6650 + ETH patch) + +Device results of the ttaw suites and probe P2 are recorded in `common/probes/p2-trace-runner.md`. In short: + +- ETH dispatch opens 12x10 with 1 and 2 CQs, WORKER 11x10, `auto` -> ETH (`tests/device/test_open_device.py`). +- In-trace `ttnn.copy` into persistent float32 / uint32 / bfloat16 tensors, ping-pong traces, RT-dev scalar params, + packed readback, 2CQ uploads (direct and staged), CQ1 readback and replay == eager are bit-exact. +- A ROW_MAJOR tensor is one page per row, and the RM `reshape` / `pad` / `concat` programs stage whole pages in L1: + `reshape_rm_program_factory.cpp` needs 2 x (destination page + 80 B) per kernel copy (two copies when the pages + divide each other) when the pages are not 16-byte aligned, out of ~1.43 MB free L1 per core + (`TT_FATAL ... RM reshape dest staging does not fit in L1: need at least 1497632 B dest + 512 B source, have + 1461248 B`, YOLOX's `[1, 1, 1, 187200]` fp32 row). Keep RM rows (last dim x element size) well below 0.5 MB; + `pack_outputs` does (section 3). +- A program-cache miss inside a capture raises with the op name and leaves the device usable. +- numpy uint32 arrays (including values >= 2**31) upload exactly; int64 arrays must not be passed to + `ttnn.from_torch` directly (they would go through bf16): use `tensors.to_host_tensor`. +- `numpy.asarray(host_buffer)` fails and `numpy.from_dlpack` is read-only; `torch.from_dlpack` is a writable alias + (`HostStaging`). +- `ttnn.matmul` with an explicit HiFi4 + fp32 config: max |error| 0.03 vs fp64 on a 64x512x128 product; LoFi 2.4. +- Runner cost (34-program toy graph, 4 MiB input): `replay` b2b equals a raw `execute_trace` (0.857 ms); streaming + `run()` per frame 1.07 ms (1CQ), 1.09 ms (direct 2CQ: CQ1 must wait for the previous replay before overwriting the + input), **0.92 ms with `stage_inputs=True`** (upload hidden behind the replay). Use staging when uploads are large + and frames are pipelined; it buys nothing for a single synchronous request. +- `ttnn.to_torch` of a host tensor copies, so `read()` results stay valid after later reads. +- Review of 0.2.0 (ttaw 0.3.0, `logs/ttaw/review_*`): `ttnn.reshape` of a ROW_MAJOR tensor returns a view on the + same buffer, and `ttnn.deallocate(view)` (default `force=True`) frees the base tensor's memory + (`review_force_dealloc_probe.log`); a program first enqueued after a capture is a "program_cache" buffer the + allocation tracker refuses before the next replay (`review_old_staged_tracking_probe.log`). Ping-pong states + written through `ctx.write_target` (`output_tensor=`), D16 bank switches between two streams, `reset_state` from + a bank, `write_input` followed by a CQ1 upload, and a 2CQ run after a partially failed capture are bit-exact on + the p150b in 1CQ and 2CQ; with `TT_METAL_TRACE_ALLOC_TRACKING=1` bank switches and the 2CQ staged path allocate + nothing after capture. +- Where the 13 specs' other needs live: fp32 row gathers are `ttnn.gather` (FLOAT32, TILE; bit-exact but ~1 us per + picked index), bf16 rows `ttnn.embedding` (3-5 us; probe P13, `probes/g1-dispatch-genericop-state.md`); top-k, + argmax, conv / pool, attention / grid_sample facts are in `probes/g2..g4-*.md`; the op wrappers C17-C23 and + kernels K1-K11 are added by the first port that needs them (PLAN.md 1.2-1.3). RT-dev values of any shape are + `add_param` (re-uploaded only when they change) or, when they change every frame (index tables), `add_input`; + shape buckets are variants sharing `add_output` buffers; trace segments around a host fallback are variants handed + off through states or a read + `run(inputs=...)`; serve profiles pin `_VARIANT` / `_NUM_CQS` / `_DISPATCH`, + checked by `server.smoke.check_pinned`; temporal state per stream is `StreamBanks` (section 3). +- Device tests: run only your own files. `common/tests/device/` also holds the probe agent's `test_probe_p*.py`, + some of which spawn device subprocesses and must not run inside another pytest process. + +--------------------------------------------------------------------------------------------------------------------- + +## 14. Image pre-processing (C14, `ttaw.image`) + +Bit-exact host emulations of the resize kernels the Autoware camera nodes deploy, driven by per-source-size lookup +tables computed once and cached (a camera keeps its size). numpy only. A standard resize in their place moves the +network input enough to fail the PCC gates (YOLOX: `cv2.resize` drops the det-output PCC to 0.94, +research/yolox/SPEC.md section 3), so each preset is tested against a scalar transliteration of the deployed code. + +```python +from .ttaw.image import yolox_letterbox, yolox_letterbox_geometry + +bgr = rgb[:, :, ::-1] # Autoware feeds BGR8 (cv_bridge toCvCopy BGR8) +x_u8, geom = yolox_letterbox(bgr) # (960, 960, 3) uint8 HWC, BGR, pad 114 +x_onnx = yolox_letterbox(bgr, layout="nchw")[0].astype("float32") # == the ONNX input `images` (1, 3, 960, 960) +geom.scale, geom.resized_hw, geom.mask_hw # float32 scale (decoder), (r_h, r_w), mask crop + +from .ttaw.image_linear import sceneseg_preprocess, sceneseg_resize, opencv_resize_nearest +u8 = sceneseg_resize(bgr) # (320, 640, 3) uint8: cv::resize INTER_LINEAR, exact +x = sceneseg_preprocess(bgr, channel_order="bgr") # == the SceneSeg ONNX input (1, 3, 320, 640) float32 +mask_src = opencv_resize_nearest(mask_320x640, bgr.shape[:2]) # the node's INTER_NEAREST back to the source size +``` + +| name | meaning | +|---|---| +| `yolox_letterbox(image, dst_hw=(960, 960), *, pad_value=114, layout="hwc", out=None)` | autoware_tensorrt_yolox `preprocess.cu:43-129` @ 9ceaccf: inverted half-magnitude "bilinear" weights, `fmaf` single roundings, float -> int truncation between the passes, `lroundf`, top-left letterbox (114), channel order kept -> `(uint8 HWC or (1, C, H, W), LetterboxGeometry)`. Needs a source of at least 2x2 | +| `yolox_letterbox_geometry(src_h, src_w, dst_h=960, dst_w=960)` -> `LetterboxGeometry` | `scale` (exact float32 value as a float; `scale_f32`), `resized_hw` = `(r_h, r_w)`, `mask_hw` (Autoware's mask crop, the same pair), `src_hw`, `dst_hw`, `pad_value`, `to_dict()` | +| `yolox_letterbox_lut(src_h, src_w, ...)` -> `YoloxLetterboxLUT` | the cached tables (`rows`, `row_w`, `cols`, `col_w`, `k`); `.apply(image, *, out=None, chunk_rows=96)` (any channel count) | +| `LUTCache(maxsize=8)` | the thread-safe LRU the presets use (`.get(key, make)`, `.hits`, `.misses`, `.clear()`) | +| `lroundf(x)`, `f32(x)` | C `lroundf` (half away from zero) and one float32 rounding, for presets and their tests | +| `bevdet_nearest_crop(image, *, channels="rgb", src_hw=(900, 1600), dst_hw=(256, 704), crop_hw=(140, 0), out=None)` | the `bevdet::Preprocess` plugin of autoware_tensorrt_bevdet (bevdet_vendor 0.2.1, research/bevdet/SPEC.md 3.2): nearest resize by r = (float)dst_w / src_w = 0.44f and crop (140, 0) of the 1600x900 image, `roundf(i / r + crop_h / r)` rows (318, 320, 323, ..., 898) and columns (0, 2, 5, ..., 1598) in float32 -> `(3, 256, 704)` uint8 **B, G, R** planes. `channels` is the input order (`"rgb"` from `ttaw.io.load_image`, `"bgr"` from OpenCV). The node's `cv::resize` of other image sizes to 1600x900 is the caller's (OpenCV) | +| `bevdet_nearest_lut(src_hw, dst_hw, crop_hw)` -> `NearestCropLUT` | the cached gather (`rows`, `cols`, `resize`); `.apply(image, *, channels, out)`, `.apply_planes(planes [..., C, H, W])` | +| `bevdet_normalize(crop, mean=BEVDET_MEAN, std=BEVDET_STD)` | `(x - mean[c]) / std[c]` in float32 on B, G, R planes (`BEVDET_MEAN` / `BEVDET_STD` are the "RGB" ImageNet numbers, applied to BGR planes as trained) | +| `bevformer_preprocess(images, *, mean=BEVFORMER_MEAN_BGR, std=BEVFORMER_STD, scale=0.8, pad_divisor=32, out=None, workers=1)` | autoware_tensorrt_bevformer's pipeline (research/bevformer/SPEC.md 3.1): N BGR uint8 images of one size (the node's 1600x900) -> normalise first (`(x - mean_c) / std_c` as OpenCV's `convertTo`: BGR means 103.53 / 116.28 / 123.675, std 1) -> `cv::resize` x0.8 INTER_LINEAR **on float** -> zero pad bottom / right to /32 -> `[N, 3, 736, 1280]` float32 B, G, R planes (the ONNX input `image` without its leading 1). `workers` threads over the cameras. The node's uint8 `cv::resize` of other sizes to 1600x900 is the caller's (OpenCV) | +| `bevformer_input_geometry(src_hw=(900, 1600), scale=0.8, pad_divisor=32)` | `{"src_hw", "resized_hw", "padded_hw"}`: `(int)(src * 0.8f)` in float32 and the /32 pad: (720, 1280) and (736, 1280) deployed (the network normalises its UV by the padded size) | +| `bevformer_normalize(image, mean, std)` | the normalisation alone, HWC float32 | +| `opencv_linear_lut(src_hw, dst_hw, *, strict=True)` -> `LinearResizeLUT`; `opencv_resize_linear_f32(image, dst_hw, *, strict=True)` | OpenCV's float INTER_LINEAR (`resize.cpp` coefficients `f = (float)((d + 0.5) * scale - 0.5)`, border clamps; each pass `fma(x1 - x0, w, x0)` in float32, horizontal first), cached tables (`rows0`, `rows1`, `wy`, `cols0`, `cols1`, `wx`); `.apply(image HWC)`, `.apply_planes(planes [..., H, W])` (strided block slices for periodic taps). `strict=True` refuses geometries not verified bit-exact (weights not multiples of 1/8: OpenCV computes other scales' coefficients differently) and exact 2x down-scales (OpenCV's INTER_AREA path) | +| `PRESETS` | name -> (function, description) of the implemented presets | +| `image_area.meteor_resize(image, *, out=None)` | METEOR (tier4/METEOR `hf/onnx_smoke_test.py:42-44`, research/meteor/SPEC.md 3): one uint8 frame (any channel order) -> 432x768 with OpenCV `INTER_AREA` (a copy when already 432x768). Down-scaling only (raises `ValueError` for any axis smaller than the target) | +| `image_area.opencv_resize_area_u8(image, dst_hw, *, out=None)`, `opencv_area_lut(src_hw, dst_hw)` -> `AreaResizeLUT` | OpenCV `cv::resize(INTER_AREA)` of uint8 HxW / HxWxC images, any down-scale: `resizeArea_` (`computeResizeAreaTab` float weights from double, horizontal then vertical float32 accumulation in table order, `cvRound`) and `resizeAreaFast_` for integer factors (float32 `s * (1.f / area)`, except the 2x2 `fast_mode` of 1-, 3- and 4-channel images: `(s + 2) >> 2`); cached tables (`fast`, `rows`, `row_w`, `cols`, `col_w`) | +| `image_area.scale_intrinsics(K, src_hw, dst_hw=(432, 768))` | K of the resized image: row 0 x dst_w / src_w, row 1 x dst_h / src_h (METEOR `extract_gt.py:128-130`; anisotropic resizes keep fx != fy), float64 | +| `image_linear.sceneseg_preprocess(image_bgr, *, channel_order="bgr")` | VisionPilot SceneSeg (`middleware_recipes/common/backends/onnx_runtime_backend.cpp:41-60` @ vision_pilot 04fa3e80; research/sceneseg/SPEC.md 3): one HxWx3 B, G, R uint8 frame -> the ONNX input `float32 [1, 3, 320, 640]`, bit-exact = `sceneseg_normalize(sceneseg_resize(image))`. `channel_order`: `"bgr"` (deployed: B, G, R planes, BGR-ordered ImageNet constants) or `"rgb"` (training order: R, G, B planes, RGB constants; a channel permutation of the bgr input) | +| `image_linear.sceneseg_resize(image_bgr, *, dst_hw=(320, 640), out=None)`, `sceneseg_normalize(resized_bgr, *, channel_order="bgr")` | the node's squash `cv::resize(img, Size(640, 320))` (INTER_LINEAR, aspect not kept) -> uint8 `(320, 640, 3)`, channel order kept (the TT device input); the normalisation as OpenCV computes it: `convertTo(CV_32F, 1/255)` (float32 product), `cv::subtract(Scalar)` in float32, `cv::divide(Scalar)` in double rounded to float32 (a float32 division differs by 1 ulp in ~27 % of the values), `cv::split` -> NCHW. `SCENESEG_INPUT_HW`, `SCENESEG_MEAN_RGB`, `SCENESEG_STD_RGB`, `SCENESEG_CHANNEL_ORDERS` | +| `image_linear.opencv_resize_linear_u8(image, dst_hw, *, out=None)`, `opencv_linear_u8_lut(src_hw, dst_hw)` -> `LinearResizeU8LUT` | OpenCV `cv::resize(INTER_LINEAR)` of uint8 HxW / HxWxC images, any scale: `resize()`'s float coefficients (`fx = (float)((dx + 0.5) * scale - 0.5)`, x-border fix only), `cvRound(w * 2048)` weights, the exact int32 horizontal pass and the uchar fixed-point vertical pass `(((b0 * (h0 >> 4)) >> 16) + ((b1 * (h1 >> 4)) >> 16) + 2) >> 2`; a copy for equal sizes and OpenCV's INTER_AREA fast path (`image_area`) for an exact 2x down-scale. Cached tables (`mode`, `cols0`, `cols1`, `a0`, `a1`, `rows`, `r0`, `r1`, `b0`, `b1`); `.apply(image, *, out=None)` | +| `image_linear.opencv_resize_nearest(image, dst_hw, *, out=None)`, `opencv_nearest_lut(src_hw, dst_hw)` -> `NearestResizeLUT` | OpenCV `cv::resize(INTER_NEAREST)` (`resizeNN`: `min(floor(x * (1.0 / inv_scale)), W - 1)` in double) of HxW / HxWxC images of any dtype (a gather): e.g. SceneSeg's 320x640 mask back to the source size (`run_model_node.cpp:117-188`) | +| `image_triangle.streampetr_preprocess(image, *, preset="autoware_main", channels="rgb", dst_hw=(480, 640), fma=False, out=None)` | autoware_camera_streampetr (universe main @ 9ceaccf, `lib/network/preprocess_kernel.cu:73-176` + `camera_data_store.cpp:287-316`; research/streampetr/SPEC.md 3.1): one uint8 `(H, W, 3)` frame (`channels` = its byte order) -> the float32 network input `(3, 480, 640)`: `resize = max(480 / H, 640 / W)` (float32), the virtual resized image `((int)(H * resize), (int)(W * resize))`, top rows cropped (bottom kept) and the width centred, a PIL-style triangle filter with adaptive support (`centre = (r + 0.5) * scale - 0.5`, taps `ceilf(centre - support) .. floorf(centre + support)`, `w = max(0, 1 - |d| / support)` per axis, `w = wx * wy`), accumulated **y outer, x inner** in float32, divided by the weight sum, then `(v - mean[c]) / std[c]`, planar. `preset` (`STREAMPETR_PRESETS`): `autoware_main` (RGB planes, RGB stats; PLAN D7 default), `autoware_0.52` (an rgb8 camera on 0.52.1: RGB planes, BGR-ordered stats), `awml_training` (BGR planes, BGR stats). `fma=True` models nvcc's default `--fmad=true` contraction (`fmaf` centre, channel and weight sums); `False` (default) rounds every C operation once, as the SPEC's port and the research goldens | +| `image_triangle.streampetr_preprocess_batch(images, *, preset, channels, dst_hw, fma, workers=1, out=None)` | the cameras of one frame -> `(N, 3, 480, 640)` float32 (each camera its own size); `workers` threads over the cameras | +| `image_triangle.autoware_resize_geometry(src_h, src_w, dst_h=480, dst_w=640)` -> `TriangleGeometry`, `triangle_resize_lut(src_hw, dst_hw=(480, 640), *, geometry=None, fma=False)` -> `TriangleResizeLUT` | `calculate_image_processing_params` in float32 (`src_hw`, `resized_hw`, `roi_hw`, `roi_start` (y, x), `resize`, `to_dict()`); the cached per-size tap tables (`rows`, `row_w`, `cols`, `col_w`; zero-weight padding) with `.apply(image u8 HxWxC) -> (C, roi_h, roi_w)` float32 weighted means (before the normalisation; any channel count). `geometry=` serves other nodes of the kernel family (BEVFusion's camera branch: verify its kernel first) | +| `image_triangle.fma_f32(a, b, c)` | `fmaf` on float32 arrays: one rounding, exact (TwoSum-corrected float64; a float64 sum landing on a float32 tie is broken towards the exact value) | + +Facts: 0 mismatches against the scalar kernel on 12 source sizes and on the research golden input of the Autoware +test image; 40-170 ms per frame on the shared host (1080p about 130 ms), a few ms per new size for the tables. The +BEVDet gather equals the research golden network input of the nuScenes key-frame `sample0` bit for bit +(`tests/host/test_image_bevdet_host.py`). The BEVFormer pipeline equals OpenCV 4.8.1 and 4.11.0 bit for bit on +arbitrary float images for every x0.8 down-scale (the textbook two-product interpolation differs in ~10 % of the +samples), the research reference's OpenCV pipeline on random uint8 images, and the research golden `image` of the +nuScenes key-frame `sample0` (`tests/host/test_image_bevformer_host.py`, which also holds a scalar exact-rational +oracle); 0.5 s per 6 x 1600x900 frame single-threaded on the shared host, 0.15 s with `workers=4` (OpenCV itself: +0.16 s). The device version (K9, uint8 upload + on-device resample) is an optimization item. Presets still to add +(their ports): BEVFusion triangle resize with truncation (PLAN.md C14; `ttaw.image_triangle` holds the same kernel +family). The METEOR `INTER_AREA` preset (`ttaw.image_area`) equals `cv2.resize` of OpenCV 4.8.1 and 4.11.0 bit +for bit at every METEOR source size tested and a scalar transliteration of `resize.cpp` on small geometries +(`tests/host/test_image_area_host.py`); about 0.1-0.15 s per 1920x1080 frame, 0.4 s at 2880x1860, numpy single thread. +The SceneSeg preset (`ttaw.image_linear`, 0.16.0) equals `cv2.resize` INTER_LINEAR / INTER_NEAREST and OpenCV's +`subtract` / `divide` scalar arithmetic of OpenCV 4.8.1 and 4.11.0 bit for bit on 23 fixed and 40 random geometries +(down / up, odd sizes, 1-4 channels, the exact-2x and copy shortcuts), a scalar transliteration of `resize.cpp`, the +23 research device inputs of the SceneSeg public-data set (`input_u8_bgr_640x320.png`) and, within one ulp, the SPEC +reference's float pre-processing (`tests/host/test_image_sceneseg_host.py`); 25-35 ms per 640x320 squash (414x727 to +1080p) and about 8 ms for the normalisation on the shared host, numpy single thread (OpenCV: 2-3 ms). +The StreamPETR preset (`ttaw.image_triangle`, 0.17.0) equals a scalar transliteration of the CUDA kernel bit for bit +in both contraction models (4 small geometries: down / up-scaling, crops, identity; 300 sampled pixels + the corners of +a 1920x1080 -> 480x640 frame), a scalar transliteration of `calculate_image_processing_params` on 15 source sizes, the +SPEC's numpy port (`research/streampetr/scripts/sp_common.py`) and the sha256 of the float32 network inputs of the +research goldens of the PandaSet ship sample (`tests/host/test_image_triangle_host.py`). The two contraction models +differ in ~43 % of the inputs by at most ~3e-5 (normalised units; the centre's last ulp moves the tap weights), far +below the device's bf16 input rounding. About 0.25-0.4 s per 1920x1080 camera on the shared host (numpy, one +thread), ~1.1 s per five-camera frame with `workers=4`. + +--------------------------------------------------------------------------------------------------------------------- + +## 15. LiDAR host pipeline (C10 `ttaw.geometry`, C11 `ttaw.pointcloud`, C12 `ttaw.voxelize`, C15 `ttaw.nms`, C16 `ttaw.decode`) + +Host pre- and post-processing of the Autoware LiDAR detectors, built by the CenterPoint port (research/centerpoint/ +SPEC.md sections 3 and 5) for CP, PP, TF and BF. numpy only. Each reproduces the deployed C++ / CUDA, including its +float32 arithmetic where that decides a cell or a threshold; the non-deterministic parts of Autoware (time-seeded +point shuffle, atomic slot / pillar order, unstable sorts) become fixed, documented policies (PLAN.md D1). + +```python +from .ttaw.pointcloud import StreamDensifiers, densify_sweeps, nonzero_rows +from .ttaw.voxelize import PillarGridSpec, assign_pillars, decorate_pillar_features, canvas_gather_index +from .ttaw.decode import CenterHeadDecodeConfig, decode_centerhead, sort_by_score, to_detected_objects +from .ttaw.nms import circle_nms, iou_bev_nms, ClassRemapper + +spec = PillarGridSpec((-76.8, -76.8, -4.0), (76.8, 76.8, 6.0), (0.32, 0.32, 10.0), 32, 40000, order="flipped_x") +cache = StreamDensifiers(num_past_frames=1).get("front") # Autoware's PointCloudDensification, per stream +cache.enqueue(xyz, stamp_s, T_world_from_ego) # False: no pose -> frame skipped, not cached +pts, info = cache.sweep_points() # (N, 4) x, y, z, time_lag; current sweep first +pil = assign_pillars(pts, spec, shuffle_seed=0) # first 32 per cell, ids in flipped-x order +feats = decorate_pillar_features(pil, spec, encoder_in_feature_size=9) # (40000, 32, 9) generateFeatures_kernel +index = canvas_gather_index(pil.coords, pil.num_pillars, spec, sentinel=40000) # NHWC canvas = table[index] +boxes = sort_by_score(decode_centerhead(heads, cfg)) # cfg = CenterHeadDecodeConfig.create(...) +keep = circle_nms(boxes.x, boxes.y, 0.5) +objs = to_detected_objects(boxes.take(keep), cfg.class_names) # ObjectClassification labels, yaw_ros +objs = objs.take(iou_bev_nms(objs.x, objs.y, objs.length, objs.width, objs.yaw, objs.label)) +objs.label = ClassRemapper.from_params(remapper_yaml_params).apply(objs.label, objs.length, objs.width) +``` + +| module / name | meaning | +|---|---| +| `geometry.compose`, `invert_rigid`, `transform_points`, `as_transform`, `transform_from_xyz_rpy`, `rotation_from_rpy` | float64 rigid transforms (`T_a_from_b`); `as_transform` accepts every `io.parse_transform` spelling | +| `geometry.invert_affine_f32`, `compose_f32`, `transform_points_f32` | Autoware's float32 path: `Eigen::Affine3f` inverse (cofactors, column-0 determinant) and product, and `generateSweepPoints_kernel` (`((m0 x + m1 y) + m2 z) + m3`, no FMA) | +| `geometry.mmdet_yaw_to_ros`, `ros_yaw_to_mmdet`, `quaternion_wxyz_from_yaw`, `yaw_from_quaternion_wxyz`, `yaw_from_rotation`, `wrap_angle` | yaw conventions (`yaw_ros = -yaw_net - pi/2`, float result as `ros_utils.cpp:55`) | +| `pointcloud.SweepDensifier(num_past_frames=1, cloud_capacity=2_000_000)` | one stream's cache: `.enqueue(xyz, stamp_s, T_world_from_sensor=None, *, intensity=None) -> bool` (pose required when `num_past_frames > 0`; a missing pose leaves the cache unchanged), `.sweep_points(*, point_feature_size=4\|5) -> (points, DensifyInfo)` (float32 `world2current @ past2world` per sweep, `time_lag = float32(t_now - t_sweep)`, capacity drops a sweep and every older one), `.reset()`, `.stamps`, `.describe()` | +| `pointcloud.StreamDensifiers(..., max_streams=16)` | per-`stream.id` caches with LRU eviction: `.get(id, reset=False)`, `.forget(id)`, `.streams`, `.describe()` | +| `pointcloud.densify_sweeps(current_xyz, [(xyz, time_lag_s, T_current_from_sweep)], ...)` | the stateless client-side variant -> `(points, DensifyInfo)` | +| `pointcloud.finite_rows`, `nonzero_rows`, `range_mask`, `seeded_permutation(n, seed)` | hygiene masks (`nonzero_rows` = `|x|+|y|+|z| > 0`, the research rule for organized-cloud fillers; `range_mask` half-open, NaN fails) and `default_rng(seed).permutation(n)` | +| `voxelize.PillarGridSpec(range_min, range_max, voxel_size, max_points_per_pillar=32, max_pillars=40000, order="flipped_x"\|"raster")` | grid sizes in float32 as `centerpoint_config.hpp` (`grid_x`, `grid_y`, `num_cells`) | +| `voxelize.assign_pillars(points, spec, *, shuffle_seed=None) -> PillarSet` | deterministic pillars: optional seeded permutation, half-open range, float32 `floor((v - min) / voxel)` clamped, first K per cell in input order, ids in ascending cell order capped at `max_pillars`. `PillarSet`: `points (P, K, F)`, `num_points`, `coords (z, y, x) = (0, iy, ix)`, `num_pillars`, `num_pillars_total`, `overflow`, `stats()` | +| `voxelize.decorate_pillar_features(pillars, spec, *, encoder_in_feature_size=9\|10\|11)` | `generateFeatures_kernel`: raw columns, offsets to the slot-ordered float32 mean, offsets to the pillar centre `voxel/2 + coord*voxel + min`; padded slots 0 | +| `voxelize.scatter_canvas(features, coords, n, spec)`, `canvas_gather_index(coords, n, spec, *, sentinel)` | `scatterFeatures_kernel` (C, H, W) and its gather form for the device (C19): cell `iy * grid_x + ix` -> pillar row or the zero sentinel row | +| `decode.CenterHeadDecodeConfig.create(class_names, voxel_size_xy, range_min_xy, downsample_factor, distance_bin_upper_limits, score_thresholds, yaw_norm_thresholds, has_variance=False, has_twist=False)` | Autoware's coercions (thresholds outside [0, 1) -> 0, ascending bins, one yaw threshold per class); thresholds `(bins, classes)`; `.with_score_threshold(v)` | +| `decode.decode_centerhead(heads, cfg) -> DecodedBoxes`, `sort_by_score` | `generateBoxes3D_kernel` in float32 (first-max label, no +0.5 offset, bins, `score < thr` drops, yaw-norm gate, dims (w, l, h), `atan2`), rows in cell order; stable descending sort (ties by cell). No max-pool, no top-K | +| `decode.head_values_at(heads, cells) -> {name: (C, K)}`, `decode_cells(values, cells, grid_w, cfg) -> DecodedCells` | the same decoder at given cells, dropping none (agreement metrics: device vs fp32 reference at the same cells). `DecodedCells`: the `DecodedBoxes` fields (bit-identical to `decode_centerhead`'s row for every cell it keeps) plus `yaw_norm`, `distance_bin`, `score_threshold`, `yaw_norm_threshold`, `passes_score`, `passes_yaw_norm`, `.keep`, `.take`, `.to_boxes()`; `head_values_at` stores the maps at a cell set (compact goldens) | +| `decode.to_detected_objects(boxes, class_names, *, has_twist=False) -> DetectedObjects` | `box3DToDetectedObject`: ObjectClassification labels (`AUTOWARE_LABELS`, `LABEL_IDS`, `semantic_label`), `yaw_ros`, (length, width, height), `orientation_availability` SIGN_UNKNOWN for car-like labels, object-frame twist; `.take`, `.boxes_xyzlwh_yaw()` (the `outputs.Detections3D` layout), `.to_records()` | +| `nms.circle_nms(x, y, dist)` | `circleNMS` on score-sorted boxes: float32 `dx^2 + dy^2 < dist^2` (strict), greedy, class-agnostic | +| `nms.iou_bev_nms(x, y, length, width, yaw, labels, *, search_distance_2d=10.0, iou_threshold=0.1, scores=None, sort=False)` | `perception_utils::IouBevNms::apply`: every earlier object (kept or not) suppresses, pedestrian-vs-other pairs skipped, `<=` search gate, break at `iou > thr`, keep iff `max_iou <= thr`; exact rotated-rectangle IoU (`iou_bev`, `clip_convex`, `polygon_area`, `bev_box_corners`) with the 1e-6 / 0.01 area guards | +| `nms.ClassRemapper.from_params(yaml_params)` / `.from_lists(allow, min, max)` | `DetectionClassRemapper`: `.apply(labels, length, width)` with `bev_area = float(l * w)`, first allowed destination in label order | + +Facts (CenterPoint port, research/centerpoint): voxel tensors bit-identical to the research reference (input order +and seed 0, including a 41,918-pillar overflow frame); decode + NMS + remapper give the same objects as two +independent research implementations on 16 head-map goldens; about 0.3 s voxelization and 30-120 ms +post-processing per 200k-point frame on the shared host (numpy). Host tests: +`tests/host/test_geometry_pointcloud_host.py`, `test_voxelize_host.py`, `test_decode_nms_host.py`. Still to add (their +ports): PointPainting 12-feature decoration and raster order use, TransFusion / BEVFusion decoders, OpenPCDet and +METEOR NMS variants (PLAN.md C12, C15, C16). + +--------------------------------------------------------------------------------------------------------------------- + +## 16. CNN ops: conv builders (C17 `ttaw.ops.conv`), up-sampling and interpolation (C18 `ttaw.ops.upsample`) + +Built by the YOLOX port (PLAN.md C17 / C18; probes P5, P14). A feature map is a `FeatureMap(tensor, batch, height, +width, channels)`: the device tensor holds N*H*W rows of C channels (ttnn's flattened `[1, 1, N*H*W, C]`, TILE or +ROW_MAJOR), plus the spatial size ttnn's layout no longer carries. Every builder prepares its device weights on its +first (eager) call and reuses them, so call each layer once before capturing (the `TraceRunner` warm-up does): a +host weight write inside a capture fails ("Writes are not supported during trace capture"). + +```python +from .ttaw.ops.conv import Conv2d, KSplitConv, ConvTranspose2d, FeatureMap, concat_channels, residual_add +from .ttaw.ops.upsample import upsample_nearest, Resize2d, interp_matrix + +conv = Conv2d(w_oihw, b, stride=2, activation="relu6") # padding "same" (k // 2); HiFi4 + fp32 + L1 acc +x = FeatureMap(ctx["image_bf16"], 1, 960, 960, 3) # or fmap_from_numpy(nchw, device) (under devrun) +y = conv(x) # FeatureMap [1, 1, 480*480, 32], DRAM, TILE +y = residual_add(conv2(y), y, "relu6") # RELU / RELU6 fused into ttnn.add (P5) +z = concat_channels([upsample_nearest(lateral, 2), d4]) # nearest x2: bit-exact; concat along C +head = Conv2d(w, b, output_dtype="float32") # fp32 logits for thresholds +ks = KSplitConv(w_full, b, parts=(128, 128, 128), activation="relu") # conv(concat) without the concat (P14) +up = ConvTranspose2d(w_iokk, b, stride=4, activation="relu") # k = 2: conv_transpose2d; k >= 3: linear + d2s (P5) +rs = Resize2d((16, 44), (64, 176), channels=256, mode="linear", align_corners=True) # FPN_LSS-style bilinear +``` + +| name | meaning | +|---|---| +| `Conv2d(weight, bias=None, *, stride=1, padding="same", dilation=1, groups=1, activation=None, precision=CONV_PRECISION, output_dtype="bfloat16", output_layout="tile", slicing="auto", output_memory_config="dram", conv_config=None, weight_terms=1, name="conv")` | `ttnn.conv2d` with weights prepared once per input spec. `__call__(fmap)` (or a raw tensor + `batch=`, `height=`, `width=`) -> `FeatureMap`. `slicing="auto"`: a DRAM input runs ttnn's DRAM-sliced path with automatic slice counts (P14: traceable; explicit counts are often rejected) and writes DRAM interleaved; an L1 input runs the L1 path; `"l1"` = `Conv2dL1FullSliceConfig`; `("height"\|"width", n)` explicit (diagnostics). 1x1 stride-1 convs are matmuls on the L1 path; `output_memory_config` (`"dram"` default, `"l1"`, None = the op's sharded output) places their output. `conv_config`: extra `ttnn.Conv2dConfig` fields (optimization knobs: `act_block_h_override`, `shard_layout`, double buffers, ...). `is_matmul`, `out_hw(h, w)`, `takes_l1_path(t)`, `prepared_specs`, `release()`, `describe()` | +| `weight_terms=2` | the weight and the bias as two bf16 terms (`hi = bf16(x)`, `lo = bf16(x - hi)`: ~16 significant bits; `host_terms()`), run as two convs into fp32 outputs, summed in fp32 with the activation fused into the add, cast to `output_dtype` (TILE outputs only): ~fp32 weights and bias for 2x the MACs and 3 more programs. **bf16 rounding of weights and biases, not of activations, dominates a deep CNN's error** (a per-channel bias error is the same at every pixel): YOLOX's mask agreement on PandaSet side cameras is 97.5 % with bf16 weights and 99.3 % with two terms (CPU emulation; the device matches it). The device's float32 weight path (`precision=":w=float32"`) is no substitute: 1.5x less error than bf16 vs 3.5x (fp32 output) for two terms; with a bf16 output two terms land within 1.05x of the output-rounding floor (`logs/yolox/m3_weight_precision_probe.log`, `m3_c17_weight_terms_device_r2.log`) | +| `CONV_PRECISION` | `"HiFi4+fp32+l1acc"`, the default of every builder. **Without packer L1 accumulation a conv whose inner dim spans several blocks keeps its partial sums in the output dtype (bf16) between blocks** (`conv2d_op_program_factory_common.cpp:175-178`): YOLOX 3x3 256 -> 256 PCC 0.99998 / max abs 0.25 vs 0.9999993 / 0.025 with it, at the same speed (`logs/ttaw/c17_c18_device_results*.json`) | +| `normalize_activation`, `unary_with_param`, `apply_activation`, `FUSED_ACTIVATIONS`, `BINARY_FUSED_ACTIVATIONS` | fused conv activations: relu, relu6, silu, gelu (exact), sigmoid; binary fused: relu, relu6. **HARDSWISH is refused** (fused, it is silently skipped: P5); apply `ttnn.hardswish` as its own op | +| `KSplitConv(weight, bias, parts, *, activation=None, output_dtype="bfloat16", accumulate_dtype="bfloat16", **conv_kwargs)` | `conv(concat(x_1..x_n))` as `sum_i conv_i(x_i)` (bias on the first part, activation fused into the last add or after it); `__call__([fmaps])`. `accumulate_dtype="float32"` keeps partials fp32. `split_k_weights(w, parts)` is the host split | +| `ConvTranspose2d(weight [Cin, Cout, k, k], bias=None, *, stride=k, activation=None, precision=CONV_PRECISION, output_dtype="bfloat16", method="auto", output_layout="tile", memory_config="dram")` | kernel == stride, no padding. `auto`: k = 2 -> `ttnn.conv_transpose2d`, k >= 3 -> `ttnn.linear` (bias + activation fused) + depth-to-space (RM reshape / permute / reshape); `linear_weights()` gives the (i, j, c)-ordered host matrix | +| `merge_sibling_convs(weights, biases) -> (w, b, sizes)`, `split_channels(fmap, sizes)` | one conv for sibling convs reading the same input (exact), split by `ttnn.slice` on C | +| `concat_channels(fmaps, *, memory_config="dram")`, `residual_add(a, b, activation=None)` | glue on maps of one spatial size | +| `FeatureMap` (`rows`, `hw`, `with_tensor(t, channels=None)`, `deallocate()`), `fmap_from_numpy(nchw, device, dtype, layout, memory_config)`, `fmap_to_numpy(fmap)`, `conv_padding`, `conv_out_hw`, `is_dram` | helpers (TILE uploads touch the device: run under `bin/devrun`) | +| `upsample_nearest(fmap, scale=2, *, output_layout="tile", memory_config="dram")` | integer-factor nearest (floor): TILE -> RM -> `[N, H, W, C]` -> `ttnn.upsample` -> `[1, 1, N*sH*H*sW*W, C]` -> TILE. Bit-exact (ONNX Resize nearest / asymmetric / floor) | +| `interp_matrix(n_in, n_out, *, mode="linear"\|"nearest", align_corners=False, scale=None)` -> `[n_out, n_in]` float64; `resize_matrices(in_hw, out_hw, ...)` | 1-D interpolation matrices equal to `F.interpolate` (torch's half-pixel rule clamped at 0, `align_corners`, the `scale_factor` rule); host tests compare them with torch at 1e-12 | +| `Resize2d(in_hw, out_hw, *, channels, batch=1, mode="linear", align_corners=False, scale=None, dtype="float32", output_dtype="bfloat16", precision="accurate", output_layout="tile")` | separable resize on the device for any sizes: W pass `[N*H, C, W] @ A_w^T` (transposes around one matmul), H pass `kron(I_N, A_h) @ [N*H, W'*C]`; fp32 operands by default (bf16 cannot hold weights like 511/1023; a device fp32 matmul is TF32-like, P12). `reference(nchw)` is its host float64 twin. `ttnn.upsample`'s own bilinear is half-pixel only (integer scales, sharded bf16 input) | + +Device results (`common/tests/device/test_conv_upsample_device.py`, every case eager x2, captured with cache misses +forbidden, replayed on a new input, and equal to an eager run after the capture bit for bit; log +`logs/yolox/m1_c17_c18_device_r1.log` / `_r2_l1acc.log`, results `logs/ttaw/c17_c18_device_results*.json`): all 54 +distinct conv shapes of YOLOX seg16 (1x1 to 9x9, stride 1 / 2, 960^2 down to 30^2, fused RELU6, fp32 head preds) +PCC >= 0.9999 (>= 0.9999992 for kxk with `CONV_PRECISION`); K-split 7 x 64 -> 128 0.99999; ConvTranspose k = 2 / 4 +0.99999; the 7 YOLOX nearest x2 shapes bit-exact (0.03-3.6 ms: the RM round trip dominates at 16 x 480^2); `Resize2d` +align_corners True / False, integer and non-integer, PCC >= 0.9999998. Host tests: `tests/host/test_conv_upsample_host.py` +(54) on the fake ttnn plus `tests/host/fake_ttnn_cnn.py` (conv2d / conv_transpose2d / upsample / matmul / argmax / +where / ... for CNN graph tests; `install(fake)` returns its undo). + +--------------------------------------------------------------------------------------------------------------------- + +## 17. Attention (C20, `ttaw.ops.attention`) + +Built by the Diffusion Planner port (PLAN.md C20; probe P7). `ttnn.transformer.scaled_dot_product_attention` has +four traps; `sdpa` is the only way the ports call it: + +1. `is_causal` defaults to True: with Sq == Sk it silently runs causal attention, with Sq != Sk it is rejected. + `sdpa` always passes `is_causal=False`. +2. The default `scale` is `1/sqrt(padded head dim)` (wrong after a 16 -> 32 head padding). `scale=` is required. +3. The default 32/32 chunks are 2.5-7x slower and inaccurate over >= 32k keys. `sdpa` passes a program config from + `chunk_config(sq, sk)` (the P7 table plus Sk-range rules) and refuses `program_config=None` once Sk >= 1,000. +4. **With a user mask the op masks the padded keys only at tile granularity**: key tiles past `ceil(Sk/32)` get + `-inf`, but the last partial tile is read from the mask, whose tile padding is 0 after a host tilize + (`sdpa/device/kernels/dataflow/reader_interleaved.cpp`, "Mask read"). Up to 31 padded keys then join the + softmax (score `q . k_pad`, value `v_pad`): rel-L2 0.18 instead of 0.025 at 321 x 321 with 89 valid keys, and + 0.06 for P7's masked 564 x 564 case. `sdpa` refuses a mask with `Sk % 32 != 0`: run masked attention on + `aligned_keys(n)` keys (the extra keys masked with `-inf` in the mask; free, the tiles are padded anyway). + Without a mask an unaligned Sk is exact (the op generates an element-wise padding mask). + +```python +from .ttaw.ops import attention as A + +q, k, v = A.split_qkv(qkv, 8) # [1, 1, S, 3*H*D] -> 3 x [1, H, S, D] (one program) +q, k, v = A.split_q_kv(q_proj, kv_proj, 8) # Q from LN(x), K|V from x, same S +qh = A.split_heads(q_cross, 8) # Q-only (cross-attention; K / V hoisted elsewhere) +mask = A.expand_key_bias(ctx["key_row"], 576) # [1,1,1,Sk] bias row (persistent input) -> [1,1,Sq,Sk] DRAM +o = A.sdpa(q, k, v, scale=1 / 32 ** 0.5, attn_mask=mask, concat_heads=True) # [1, 1, Sq, H*D] +runner.run("plan", inputs={"key_row": A.key_bias_row(np.r_[valid, np.zeros(12, bool)])}) # 564 real + 12 masked +wq, bq = A.pad_head_columns(w_q, b_q, 8, 16) # head dim 16 -> 32 in the weights (exact, free) +wo = A.pad_head_rows(w_out, 8, 16) +``` + +| name | meaning | +|---|---| +| `sdpa(q, k, v, *, scale, attn_mask=None, program_config="auto", compute_kernel_config=None, fp32_acc=False, concat_heads=False, memory_config=None)` | `[B, H, Sq, D]` x `[B, Hkv, Sk, D]` (bf16 / bfp8 / bfp4 TILE interleaved, D % 32 == 0) -> `[B, H, Sq, D]`, or `[B, 1, Sq, H*D]` with `concat_heads=True` (the op's `output_concat_heads`, no extra program). Mask: additive bf16 TILE in DRAM, `[1\|B, 1\|H, Sq, Sk]`, Sk tile-aligned; `-inf` is safe | +| `chunk_config(sq, sk) -> (q_chunk, k_chunk)`, `CHUNK_TABLE` | exact measured shapes (evidence string per entry), else Sk >= 16,384 q64 k1024; >= 4,096 q64 k512; >= 1,000 q64 k128; shorter q64 k64 (32 when the sequence fits one tile) | +| `sdpa_program_config(device, sq, sk, *, q_chunk=None, k_chunk=None, exp_approx_mode=None)`, `sdpa_compute_config(*, fp32_acc=False, fidelity="HiFi2", approx=True)` | `SDPAProgramConfig` over the grid read from the device; the explicit compute config equal to the op default (P7: HiFi4 / exact exp do not help; `fp32_acc` halves the max error at ~1.4x the time) | +| `split_heads(x, num_heads)`, `split_qkv(qkv, num_heads, *, num_kv_heads=None)`, `split_q_kv(q, kv, num_heads, *, num_kv_heads=None)`, `merge_heads(x)` | `[B, 1, S, H*D]` <-> `[B, H, S, D]` through `ttnn.experimental.nlp_create_qkv_heads` / `nlp_concat_heads` (one program each, bit-exact moves, the logical S kept). `split_q_kv` needs the same S for Q and K / V; cross-attention splits Q and K / V separately | +| `key_bias_row(valid) -> [1, 1, 1, Sk]`, `key_bias(valid, sq) -> [1, 1, Sq, Sk]`, `expand_key_bias(row, sq)`, `aligned_keys(n)` | host masks (0 / `-inf`, float32; upload bf16 TILE) and the traced expansion of a persistent bias row (`ttnn.repeat`, DRAM; one program per plan, not per call) | +| `padded_head_dim(d)`, `pad_head_dim(x)`, `pad_head_columns(w, b, H, D)`, `pad_head_rows(w, H, D)` | head dim padding to a multiple of 32 in the projection weights (exact: the padded output columns are 0, P7) | +| `attention_reference(q, k, v, *, scale, bias=None)`, `attention_matmul(q, k, v, *, scale, attn_mask=None, compute_kernel_config=None)` | float64 numpy oracle; small-MHA fallback with two matmuls + softmax for what SDPA rejects (e.g. fp32 Q / K / V; tile-aligned Sk) | + +Device results (`common/tests/device/test_attention_device.py`, 27 cases through `TraceRunner`: eager x2 +bit-identical, strict capture, replay on a new input set equal to an eager run bit for bit; set A nearly uniform +attention, set B peaked; gates frozen in `test_attention_device.gates.json`; `logs/diffusion-planner/c20_device_r2.log`, +results `logs/ttaw/attention_results.json`): PCC vs float64 0.99966-0.99983 at the DP shapes (576 x 576 and 352 x 352 +masked, 352 / 321 x 564), TF self 500 x 500 head dim 16 -> 32, SP self 900 x 1,668 and BF cross 500 x 32,400 +(q64 k1024, 0.46 ms); rel-L2 0.024-0.025 everywhere; split / merge / mask expansion bit-exact; the fp32 matmul +fallback PCC 0.999998. Trace ms per call: DP fusion 576 x 576 masked 0.061-0.092, DiT self 352 x 352 masked 0.063, +DiT cross 352 x 564 0.033. A canary (`test_raw_op_mask_padding_leak`) keeps the trap-4 evidence (raw op, 321 keys: +rel-L2 0.183); when tt-metal masks partial tiles it fails and the refusal can be relaxed. Watcher bring-up of the +smallest shape (`TT_METAL_WATCHER=2`) clean: `logs/diffusion-planner/c20_watcher_smoke.log`. Host tests: +`tests/host/test_attention_host.py` (27) with `tests/host/fake_ttnn_attention.py` (an SDPA fake that keeps the op's +defaults, so a caller forgetting `is_causal=False` / `scale` fails; head ops, repeat, matmul, softmax). + +--------------------------------------------------------------------------------------------------------------------- + +## 18. LiDAR device modules: gather-form scatter (C19 `ttaw.ops.gather`), SECOND + SECONDFPN (C24 `ttaw.models.second`), CenterHead (C25 `ttaw.models.centerhead`) + +Built by the CenterPoint port (PLAN.md C19, C24, C25; probes P4, P5, P14) on the C17 builders of section 16, for +CP, PP, TF and BF. Every module is built from BN-folded host (numpy) weights, prepares its device weights on its first +(eager) call, keeps every output in DRAM (interleaved TILE; L1 residency is optimization work) and traces. + +```python +from .ttaw.ops.gather import check_index, scatter_rows, gather_rows, scatter_rows_numpy +from .ttaw.models.second import ConvSpec, Second, SecondFPN, RowLinear, free_maps +from .ttaw.models.centerhead import CenterHead, merge_heads +from .ttaw.ops.conv import FeatureMap + +index = check_index(prep.canvas_index, num_rows=prep.num_pillars, sentinel=40000) # host, [1, H*W] uint32 +backbone = Second([[ConvSpec(w, b, stride=s) for w, b, s in block] for block in blocks]) # 3x3 + ReLU +neck = SecondFPN([ConvSpec(w0, b0), ConvSpec(w1, b1, stride=2, kind="conv_transpose"), + ConvSpec(w2, b2, stride=4, kind="conv_transpose")]) # 1x1, k2 ConvT, k4 ConvT +head = CenterHead(shared_spec, [128, 128, 128], [(name, hidden_spec, final_spec), ...]) # K-split + merged heads + +def forward(ctx): # a TraceRunner variant + canvas, table = scatter_rows(pillar_rows, ctx["canvas_index"]) # [1,1,P,32] rows -> [1,1,H*W,32] TILE + blocks = backbone(FeatureMap(canvas, 1, 480, 480, 32)) + return head(neck(blocks)).tensor # one fp32 [1, 1, H*W, 32] tensor (15 channels used) +maps = head.unpack(runner("default", inputs=...), 480, 480) # host: {name: (C, H, W)} +``` + +| name | meaning | +|---|---| +| `gather.check_index(index, *, num_rows, sentinel=None)` | host validation of a gather index -> contiguous `uint32 [1, N]`; every entry in `[0, num_rows)` or equal to `sentinel`, else `ValueError` (an out-of-range index reads outside the table on the device) | +| `gather.gather_rows(index, table, *, sentinel=None, layout="tile", memory_config=None)` | `ttnn.embedding` of a UINT32 `[1, N]` index into a bf16 ROW_MAJOR `[1, 1, V, D]` table -> `[1, 1, N, D]`; with `sentinel` it is P4's fast `EmbeddingsType.PADDED` form (`padding_idx=sentinel`), which **returns the table's own row there**: the row must hold zeros | +| `gather.with_zero_rows(rows, count=1)` | `[1, 1, P, D]` rows (TILE or ROW_MAJOR) -> ROW_MAJOR table `[1, 1, P + count, D]` with zero rows appended (one `ttnn.pad`) | +| `gather.scatter_rows(rows, index, *, sentinel=None, layout="tile", memory_config=None) -> (canvas, table)` | the gather-form scatter: `canvas[n] = rows[index[n]]`, zeros where `index[n] == sentinel` (default `P`); static shapes, no canvas clear; P4: 0.21 ms at CP 480 x 480 | +| `gather.gather_rows_numpy`, `gather.scatter_rows_numpy(rows, index, *, sentinel)` | host oracles | +| `second.ConvSpec(weight, bias, stride=1, padding=None, relu=True, kind="conv"\|"conv_transpose", name="")` | one BN-folded layer: conv weight `[Cout, Cin, k, k]` (padding None = `k // 2`), ConvTranspose weight `[Cin, Cout, s, s]` (kernel == stride) | +| `second.Second(blocks, *, policy=None, prefix="backbone")` | SECOND: `blocks[i]` = list of 3x3 `ConvSpec` (block stride on the first); `__call__(fmap) -> [block outputs]` (input kept, inner layer outputs freed), `out_shapes(h, w)`, `release()`, `describe()`. C17 `Conv2d` layers: a DRAM input runs ttnn's automatic DRAM slicing (P14) | +| `second.SecondFPN(deblocks, *, policy=None, prefix="neck")` | one deblock per block output: `"conv"` 1x1 -> `RowLinear`; ConvTranspose k = 1 -> `RowLinear`, k = 2 -> `ttnn.conv_transpose2d`, k >= 3 -> linear + depth-to-space (C17 `ConvTranspose2d`, P5); every output moved to DRAM; the concat is never built | +| `second.RowLinear(weight, bias=None, *, activation=None\|"relu", precision=DEFAULT_PRECISION, output_dtype="bfloat16")` | a 1x1 conv / per-row dense layer: `ttnn.linear` on `[..., R, K]` TILE rows with an explicit compute config and `core_grid` from the device (keeps the ReLU fused), output DRAM; `weight` `[K, N]` or `[N, K, 1, 1]`; device weights uploaded on the first call | +| `second.DEFAULT_PRECISION` | C17's `CONV_PRECISION` (`HiFi4+fp32+l1acc`); a `ttaw.precision.PrecisionPolicy` overrides per module (`backbone.block.conv`, `neck.deblock`, `head.shared`, `head.hidden`, `head.out`); a precision's `activations` dtype (`:a=fp32`) is the dtype of that module's OUTPUT (bf16 by default; the K-split partial sums follow the shared conv's) | +| `second.free_maps(*maps)` | deallocate feature maps / tensors that still hold their buffer (`force=False`: views are left alone, P2) | +| `second.second_numpy(blocks, x_nchw)`, `second.second_fpn_numpy(deblocks, blocks_nchw)` | fp32 torch oracles (NCHW) | +| `centerhead.CenterHead(shared, split, heads, *, policy=None, prefix="head", output_dtype="float32", k_split=True, accumulate_dtype="bfloat16")` | `relu(conv3x3(concat(inputs)))` as a C17 `KSplitConv` over `split` (P14: 3.6 vs 14.8 ms at CP 480 x 480, more accurate), then the merged heads: `relu(x @ [Wh_1 \| ... \| Wh_n] + bh)` and the block-diagonal `@ Wo + bo` (two `RowLinear`s) -> ONE `[1, 1, H*W, C_pad]` tensor (`C_pad` = sum of head channels rounded up to 32). `shared_forward(inputs)`, `heads_forward(shared)`, `__call__(inputs)`, `unpack(rows, h, w) -> {name: (C, H, W)}`, `slices`, `release()`, `describe()` | +| `centerhead.merge_heads(heads, *, pad_to=32)` | host: `(Wh, bh, Wo, bo, slices)` of `[(name, hidden 1x1 + ReLU, final 1x1)]` (exact rewrite) | +| `centerhead.centerhead_numpy(shared, heads, inputs_nchw)` | fp32 torch oracle of the literal graph (concat, conv, n heads) | + +Device results (`common/tests/device/test_lidar_models_device.py`, TraceRunner: eager x2, strict capture, replay on a +new input, replay == eager bit for bit): the CP scatter at its real shape (40,000 rows into 480 x 480) bit-exact, +sentinel rows zero; CP-shaped SECOND (4 / 6 / 6 convs, 64 / 128 / 256 channels) + SECONDFPN + CenterHead with random +weights at 64 x 64 and 128 x 128 (first stride 1 and 2) PCC >= 0.999 for every block, deblock, the shared conv and +every head vs the torch fp32 oracle (`logs/ttaw/c19_c24_c25_device_results.json`). The CenterPoint bundle runs the same +modules with the real weights (`bundles/centerpoint-p150/PORT_LOG.md`). Host tests: +`tests/host/test_lidar_models_host.py` (13) on the fake ttnn plus `tests/host/fake_ttnn_lidar.py` (`embedding` with +P4's PADDED semantics, `max`, `EmbeddingsType`; `install(fake)` returns its undo). + +--------------------------------------------------------------------------------------------------------------------- + +## 19. grid_sample helpers (C23, `ttaw.ops.deform`) + +Built by the BEVDet port (PLAN.md C23; probe P9) for AlignBEV (8 history slots warped by per-frame affine grids) +and FPN_LSS's align_corners bilinear up-sampling; BEVFormer (DCN / rotate) and METEOR extend it. Grids are built on +the host (numpy, float32) and uploaded as ROW_MAJOR fp32 tensors (a per-frame RT-dev value or a static input); +`grid_sample` is the device call with the P9 rules checked first. + +```python +from .ttaw.ops import deform + +ring = runner.add_state("ring", shape=(8, 128, 128, 96), dtype="bfloat16", layout=ttnn.ROW_MAJOR_LAYOUT) # C % 32 +grid = deform.affine_grid(transforms_8x6, (128, 128), valid=[h >= 0 for h in history]) # [8, 128, 128, 2] fp32 +runner.add_param("grid", grid, shape=grid.shape, dtype="float32", layout=ttnn.ROW_MAJOR_LAYOUT) +aligned = deform.grid_sample(ctx["ring"], ctx["grid"]) # in a variant: bilinear, zeros, align_corners=False +up = deform.grid_sample(f2_rm, static_grid) # static_grid = deform.resize_grid((16, 16), (64, 64)) +``` + +| name | meaning | +|---|---| +| `pixel_to_grid(px, size)` | pixel coordinate (0 = first pixel centre) -> normalised `(2 px + 1) / size - 1` (float32): the align_corners=False form, exact for integer `px` when `size` is a power of two | +| `kernel_pixels(g, size)` | what the device kernel recovers: `g * size/2 + (size - 1)/2` in float32 (`grid_sample_reader_common.hpp`) | +| `affine_pixel_coords(transforms, out_hw)` | `ix = a w + b h + c`, `iy = d w + e h + f` per batch (float32, one rounding per op): the BEVDet AlignBEV pixel affine | +| `affine_grid(transforms [N, 6], in_hw, out_hw=None, *, valid=None)` | fp32 grid `[N, H, W, 2]`; `valid[n] == False` -> `OUTSIDE` (an exact zero sample: an empty history slot) | +| `resize_pixel_coords(in, out, align_corners=True)`, `resize_grid(in_hw, out_hw, *, align_corners=True, batch=1)` | static bilinear-resize grids (float64 source pixels, align_corners=False normalisation); for align_corners=True every source pixel is inside, so it equals `F.interpolate` up to the weight truncation | +| `padded_channels(c)`, `pad_channels(x)` | the input channel rule (multiple of 32; pad channels stay exactly 0 through the op) | +| `emulate_grid_sample(x, grid)` | host emulation of the kernel (fp32 fractions, weights **truncated** to bf16, taps outside skipped): the device test's oracle | +| `grid_sample(x, grid, *, compute_kernel_config=None, memory_config=None, batch_output_channels=False)` | `ttnn.grid_sample` bilinear / zeros / align_corners=False; refuses a TILE input or grid, input C % 32 != 0, a grid batch != input batch, a bf16 grid, and a FLOAT32 input with `fp32_dest_acc_en` (P9 defect); output DRAM by default | +| `OUTSIDE` | `-4.0`: a normalised coordinate whose four taps miss every input | + +Rules (P9): fp32 grids for every data-dependent grid (bf16 grids: PCC 0.9988 at 53 px, 0.992 at 128 px); C padded to +a multiple of 32; grids expanded per batch (no broadcast); ROW_MAJOR interleaved / height-sharded input and grid; +never a FLOAT32 input with `fp32_dest_acc_en`; call the op with align_corners=False on the align_corners=False +normalisation of the pixel coordinates you mean (zero padding acts in pixel space, so it is the same sampling as the +model's align_corners=True formulation, but integer coordinates stay exact: an identity warp is a bit-exact copy; +the literal ac1 normalisation leaves 2.8-7.6 % of values off by up to 1/32). The bilinear weights are truncated to +bf16 (0.33-0.6 % rrmse floor; K5 removes it); C18 `Resize2d` (fp32 interpolation matmuls) is the accurate +alternative for static resizes. + +Device results (`tests/device/test_deform_device.py`, TraceRunner: eager warm-up, strict capture, replay on new +data, replay == eager bit for bit; `logs/ttaw/deform_device_results.json`): AlignBEV [8, 128, 128, 96] identity slots +bit-exact, `OUTSIDE` slots exactly 0, pad channels 0, rigid-motion slots PCC >= 0.9999 vs torch fp32 and within 2 +bf16 ulps of `emulate_grid_sample`; the FPN_LSS resizes 16 -> 64 (640 ch) and 64 -> 128 (512 ch) PCC >= 0.99999 vs +`F.interpolate(align_corners=True)`, corners exact. Host tests: `tests/host/test_deform_host.py` (13). The BEVDet +bundle runs the same calls at its real shapes (`bundles/bevdet-p150/PORT_LOG.md`: AlignBEV moving slots PCC 0.999995 +vs the CPU reference). + +--------------------------------------------------------------------------------------------------------------------- + +## 20. ResNet builders (C27, `ttaw.models.resnet`) + +Built by the BEVDet port (PLAN.md C27) for its ResNet-50 image backbone (6 cameras as batch) and CustomResNet BEV +encoder; BEVFormer and METEOR reuse it. Built on the C17 builders of section 16 from BN-folded host weights; every op +traces, each conv prepares its device weights on its first (eager) call, every activation is a DRAM-interleaved TILE +`[1, 1, N*H*W, C]` map and intermediates are freed as soon as they are consumed. + +```python +from .ttaw.models import resnet as rn + +prov = lambda module: rn.ConvParams(w[module], b[module], stride, padding_tblr, dilation, groups) # BN folded +stem = rn.ResNetStem(prov, "img_backbone.conv1") # 7x7 s2 + ReLU + maxpool 3x3 s2 +r50 = rn.ResNetStages(prov, rn.mmdet_names("img_backbone"), [3, 4, 6, 3]) # bottleneck, pytorch style +c4, c5 = r50(stem(img_fmap), keep=(2, 3), free_input=True) # FPN inputs; the rest freed +bev = rn.ResNetStages(prov, rn.custom_resnet_names("img_bev_encoder_backbone"), [2, 2, 2], block="basic", + factory=rn.conv_factory("HiFi4+fp32+l1acc:w=fp32")) +f0, f1, f2 = bev(bev_in, keep=(0, 1, 2)) +``` + +| name | meaning | +|---|---| +| `ConvParams(weight [Cout, Cin/g, kh, kw], bias=None, stride=(1, 1), padding=(t, b, l, r), dilation=(1, 1), groups=1)` | one BN-folded conv as a provider returns it (numpy float32); a provider is `module name -> ConvParams` (BEVDet's reads `OnnxWeights` by consuming node, so no kernel size or stride is re-typed) | +| `conv_factory(precision=DEFAULT_PRECISION, **conv_kwargs)` | `(module, params, activation) -> callable`: one C17 `Conv2d` per module; `precision` a spec / `Precision` or a `PrecisionPolicy` resolved per module name. **DCN hook:** pass your own factory and return a deformable layer for the modules you name | +| `mmdet_names(prefix)`, `custom_resnet_names(prefix)` | `(stage, block) -> {conv1, conv2[, conv3], downsample, block}`: `layer{s+1}.{b}.conv{k}` / `.downsample.0`, or `layers.{s}.{b}.conv{k}` / `.downsample` | +| `Bottleneck(provider, names, *, has_downsample, factory)` | `relu(conv3(relu(conv2(relu(conv1(x))))) + idt)`, stride on `conv2` (pytorch style) | +| `BasicBlock(provider, names, *, has_downsample, factory)` | `relu(conv2(relu(conv1(x))) + idt)`; the projection may be any conv (CustomResNet: 3x3 stride 2) | +| `ResNetStem(provider, conv_module, *, factory=None, pool=True)` | `maxpool3x3s2p1(relu(conv1(x)))`; the pool runs DRAM-sliced and writes TILE | +| `ResNetStages(provider, names, layers, *, block="bottleneck"\|"basic", factory=None, downsample_first=())` | `__call__(x, keep=(), *, free_input=False) -> [kept stage outputs]` (input kept unless `free_input`; every other block output freed once consumed); `stages`, `convs()` | +| `maxpool_out_hw(h, w)`, `DEFAULT_PRECISION` (= C17 `CONV_PRECISION`) | helpers | +| `resnet_numpy(x_nchw, provider, names, layers, *, block, stem=None, keep=(), downsample_first=())` | fp32 torch oracle of the same blocks | + +Precision: the ports pick it. BEVDet runs `HiFi4+fp32+l1acc:w=fp32` (fp32 weights and biases: the bf16 rounding of +the BN-folded biases, applied per channel at every pixel, cost it depth PCC 0.9998 vs 0.99993 and 3 of the 52 sample0 +boxes, at no measured time cost); C17's `weight_terms=2` is the other accurate option. Device results +(`tests/device/test_resnet_device.py`, TraceRunner protocol; `logs/ttaw/resnet_device_results.json`): stem + two +bottleneck stages on [6, 3, 128, 352] and a CustomResNet stage (3x3 s2 projection) on [1, 864, 128, 128] -> 160, +random weights, PCC >= 0.999 vs `resnet_numpy` with both precisions, replay == eager bit for bit. Host tests: +`tests/host/test_resnet_host.py` (6, fake ttnn + `fake_ttnn_cnn` + a local max-pool fake). The BEVDet bundle runs the +real ResNet-50 / CustomResNet (`bundles/bevdet-p150/PORT_LOG.md`: depth / feat PCC 0.99994, chained heads >= 0.9995 +vs the fp32 CPU reference). + +--------------------------------------------------------------------------------------------------------------------- + +## 21. Query heads: top-k (C21 `ttaw.ops.topk`), heatmap peaks (C22 `ttaw.ops.heatmap`), TransFusion head (C26 `ttaw.models.transfusion_head`) + +Built by the TransFusion port (PLAN.md C21, C22, C26; probes P6, P7, P8, P13) for TF, BF and PT: everything after the +dense heatmap -- the 3x3 local maximum, the top-K proposal selection and the query decoder -- as one traceable device +graph with one packed output. numpy only at import. + +```python +from .ttaw.ops.heatmap import LocalMax, sigmoid_heat +from .ttaw.ops.topk import TopKSelect, selection_metrics +from .ttaw.models.transfusion_head import QueryHeadWeights, TransFusionQueryHead, assemble_transfusion + +local_max = LocalMax(192, 192, 5, rule="transfusion") # "bevfusion" / "ptv3": pooled_classes, passing cells +head = TransFusionQueryHead(QueryHeadWeights(...), prefix="decoder") # host numpy weights, BN folded +head.prepare(device) # tables + LayerNorm rows (before any capture) + +def forward(ctx): # a TraceRunner variant + heat, _ = sigmoid_heat(logits_fp32) # fp32 sigmoid -> bf16 (max_pool2d / top-k are bf16) + heat_nms = local_max(heat) # [1, 1, H*W, C] TILE + out = head(lidar_feat, heat_nms, ctx.get("topk_indices")) # None: device top-k; else teacher forcing + return pack_outputs(out) # heads (K x 32 fp32), query heat (Kp x 32), indices +raw = runner("default", inputs=...) +outs = assemble_transfusion(head.unpack_heads(raw["heads"]), raw["query_heat"], raw["indices"], bev_pos, + num_classes=5, num_proposals=500, cells=36864) # cls_score0 / bbox_pred0 / dir_cls_pred0 +``` + +| name | meaning | +|---|---| +| `heatmap.local_max_masks(h, w, classes, *, rule="transfusion"\|"bevfusion"\|"ptv3", pooled_classes=None, kernel=3)` | the constant masks `(m_pass, m_eq)` (NHWC rows `[h*w, classes]`, 0 / 1, disjoint) of `keep = m_pass + (heat == maxpool_kxk_s1_p(k//2)(heat)) * m_eq`: TF interior `m_eq`, no pass; BF pooled classes (default 0-3) as TF, the others pass everywhere; PT `m_pass = 1 - m_eq` (border and unpooled classes pass) | +| `heatmap.LocalMax(h, w, classes, *, rule, pooled_classes=None, kernel=3, memory_config="dram")` | the device chain on a bf16 TILE `[1, 1, h*w, classes]` heat (`max_pool2d` p1, `eq`, `multiply` by `m_eq` or `addcmul(m_pass, eq, m_eq)`, `multiply` by the heat); masks uploaded by the first call; `.numpy(heat_rows)` its host twin; refuses fp32 (max_pool2d is bf16 only) | +| `heatmap.sigmoid_heat(logits, *, dtype="bfloat16") -> (heat, sig)` | the sigmoid in the logits' dtype (fp32 logits: fp32 sigmoid), cast once; `sigmoid_heat_numpy` its host model | +| `heatmap.local_max_numpy`, `local_max_literal_numpy(heat_nchw, rule)` | oracles: the masked chain, and each model's literal `p0` + border / unpooled-class form | +| `topk.TopKSelect(cells, classes, k)` | `__call__(heat_nhwc) -> (indices, pos, cls)` (uint32 ROW_MAJOR `[1, padded_k(k)]`): `class_major_row` (transpose + untilize + reshape to one bf16 row, index `c * cells + p` = the ONNX `TopK` order), `topk_indices` (`topk_large_indices`, descending, ties unspecified), `decode` | +| `topk.decode_class_major(indices, cells, classes, *, clamp=True)` | exact decode with fp32 comparisons (`cls = sum_c idx >= c * cells`, `pos = idx - cls * cells`; no division), clamped so a 0xFFFFFFFF sentinel cannot become an out-of-range gather index; uint32 ROW_MAJOR rows for `ttnn.embedding` | +| `topk.padded_k(k)`, `K_MULTIPLE`, `MASK_VALUE` (-1e30), `SENTINEL` | k rounded up to a multiple of 16 (500 -> 512); mask scores with a finite value, never `-inf` (P8) | +| `topk.topk_numpy(scores, k)`, `class_major_numpy`, `decode_class_major_numpy`, `selection_metrics(dev_idx, ref_idx, *, reference_scores=None, flat_scores=None, strong=0.1)` | ONNX-semantics host top-k (ties to the lower index), layouts, and the set-based metrics of a device selection (`shared`, `strong_total` / `strong_kept` / `strong_overlap` / `strong_missed`) | +| `topk.topk_with_values(scores_tile, k)` | the bf16 `ttnn.topk` composite (values + indices, P8) | +| `transfusion_head.QueryHeadWeights(height, width, num_classes, num_proposals, class_table, bev_pos, self_posembed, cross_posembed, self_attn, cross_attn, norms, ffn, heads)` | model-agnostic host record (`MHAWeights`, `PosEmbedMLP`, `PredictionHead`; `[in, out]` layouts): `.validate()`, `.query_pos_table()` / `.key_pos_table()` (the position MLPs on `bev_pos`, float64), `.merged_heads()` | +| `transfusion_head.TransFusionQueryHead(weights, *, policy=None, prefix="head", attn_fp32_acc=False)` | `__call__(lidar_feat, heat_nms, indices=None, taps=None) -> {"heads", "query_heat", "indices"}`; pieces `init_queries`, `keys`, `decoder`, `predict`, `query_heat`; `unpack_heads(raw) -> {name: (c, K)}`; `prepare(device)`, `release()`, `describe()` | +| `transfusion_head.QueryHeadOracle(weights)` | the literal head in torch fp32 (posembed MLPs on `bev_pos[pos]`, unpadded attention, six heads) from NCHW maps and given indices: the device tests' reference | +| `transfusion_head.assemble_transfusion(heads, query_heat, indices, bev_pos, *, num_classes, num_proposals, cells)` | TF's (mmdeploy-patched) outputs on the host in float32: `cls_score0 = query_heat * sigmoid(heatmap)`, `bbox_pred0 = [center + query_pos, height, dim, vel]`, `dir_cls_pred0 = rot` | + +Device form of the head (every rewrite exact in real arithmetic; the oracle runs the literal forms): queries +`lidar_feat[pos] + class_table[cls]` (bf16 `ttnn.embedding` row gathers, uint32 indices from the exact decode); the +query position embedding is a function of `pos` only, so it is a constant **QPE table** `self_posembed(bev_pos)` +gathered by `pos` (`query_pos` itself, `k + 0.5` up to 191.5, never passes through bf16: S:transfusion:241); keys +`lidar_feat + KPE` with the constant `KPE = cross_posembed(bev_pos)`; head dim 16 -> 32 in the projection weights +(C20 `pad_head_columns` / `pad_head_rows`, exact), `scale` from the weights, C20 `sdpa` (non-causal, chunk table); +self-attention on the logical K queries with no mask (Q / K / V from one fused projection); LayerNorm with the +model's epsilon (`ttnn.layer_norm(o, residual_input_tensor=x, epsilon=...)`); merged heads (`relu(x @ [Wh_i]) @ +blockdiag(Wo_i)`, fp32 out). One decoder layer only (a second layer would re-embed the predicted centres: no constant +table). Precision `HiFi4+fp32:w=fp32` (`default_policy()`; per module `.self_attn` / `.cross_attn` / `.ffn` / +`.heads` / `.norm`). + +Device results (`tests/device/test_transfusion_head_device.py`, 13 cases through `TraceRunner`: eager x2, strict +capture, replay on a new input set, replay == eager bit for bit; watcher bring-up of the small shapes clean; +`logs/transfusion/m1/j2_ttaw_{small_watcher,full}.log`, results `logs/ttaw/c21_c22_c26_device_results.json`, gates +frozen in `test_transfusion_head_device.gates.json`): the masked local max is **bit-exact** against each model's +literal form at TF 192x192x5, BF 180x180x5 and PT 256x256x7 (plateaus and bright borders: ties everywhere); the +top-512 of TF 184,320 / BF 162,000 / PT 458,752 values has exactly the host's value multiset, unique in-range +indices in descending order, and an exact decode (also exhaustively over all 184,320 TF indices); surviving `-inf` +scores give 0xFFFFFFFF (22 of 32 in the canary) and the decode clamps them in range; the query head with random +weights, teacher-forced proposals, vs the literal fp32 oracle: TF (36,864 keys) PCC >= 0.9970 (queries 0.999998, +decoder layers >= 0.9996, heads >= 0.9970), BF (32,400 keys, the (i, j) `bev_pos`) >= 0.9922, query heat exact. +The TransFusion bundle runs the same modules with the deployed weights (`bundles/transfusion-p150/PORT_LOG.md`). +The selection chain at TF size (sigmoid, local max, class-major row, top-512, decode, one row gather) replays in +about 1.1 ms (`logs/transfusion/m1/exp1_full_*.log`: 2.0-2.4 ms for two chains). Host tests: +`tests/host/test_transfusion_head_host.py` (29) with `tests/host/fake_ttnn_query.py` (`topk_large_indices` with +HIGHEST-index ties, bf16-only `max_pool2d`, `layer_norm` with ttnn's 1e-12 default epsilon, `eq` / `ge` / +`subtract` / `addcmul`, an `embedding` that also records computed indices inside a capture). + +--------------------------------------------------------------------------------------------------------------------- + +## 22. Pillar feature net and input staging (C24 companion, `ttaw.models.pillars`) + +The PillarFeatureNet of the pillar detectors (CenterPoint, PointPainting, TransFusion) on the device and the host +staging of its two inputs. Promoted by the PointPainting port (ttaw 0.19.0) from the CenterPoint bundle's +`tt/pfn.py` + `FrameStaging` (device-verified there: PFN PCC 0.99997) with the TransFusion bundle's generalisation +(any hidden / output width; pillars of K < 32 points). The CenterPoint and TransFusion bundles keep their own copies +until they re-vendor and switch (their PORT_LOGs note it). + +```python +from .ttaw.models.pillars import FrameStaging, PfnPlan, TtPillarFeatureNet, features_shape, pack_inputs +from .ttaw.ops.gather import scatter_rows + +plan = PfnPlan.from_weights(w0, b0, w1, b1) # BN-folded [in, out] matrices: (F, H), (2H, O) +pfn = TtPillarFeatureNet(plan, num_pillars=40000, precision="HiFi4+fp32:w=fp32:a=fp32") # a= : hidden h / z dtype +runner.add_input("features", shape=features_shape(40000), dtype="bfloat16") # ROW_MAJOR, 1 KB / pillar +runner.add_input("canvas_index", init=np.full((1, H * W), 40000, np.uint32), dtype="uint32") + +def forward(ctx): # a TraceRunner variant + rows = pfn(ctx["features"]) # [1, 1, P, O] bf16 ROW_MAJOR + canvas, table = scatter_rows(rows, ctx["canvas_index"], sentinel=40000) # C19: NHWC [1, 1, H*W, O] TILE + ... +staging = FrameStaging(40000, H * W) # persistent host tensors, written in place per frame +out = runner("default", inputs=staging.write(features, canvas_index, num_pillars)) +``` + +| name | meaning | +|---|---| +| `PfnPlan.from_weights(w0, b0, w1, b1, *, in_padded=16)` | the device matrices of the identity-folded split PFN: `W0 (16, H_pad)`, `Wz = [I_H \| W1a \| 0] (H_pad, Z)`, `Wy = [W1b ; I_O ; 0] (Z, O)` (`H_pad` = H rounded up to 32, `Z` = H + O rounded up to 32); F <= 16, O % 32 == 0 (the C19 gather table). `.forward_numpy(features, *, hidden_round=None)`: its float32 host emulation (`hidden_round` models the hidden dtype) | +| `pfn_reference_numpy(features, w0, b0, w1, b1)` | the literal encoder in float64 (concat form; max over the given slots, padded slots included): the oracle | +| `TtPillarFeatureNet(plan, *, num_pillars, precision="HiFi4+fp32:w=fp32", name="pfn")` | `[1, 1, P, 32*16]` bf16 ROW_MAJOR -> `[1, 1, P, O]` bf16 ROW_MAJOR rows: reshape to `[1, 1, P*32, 16]`, tilize, `linear` W0 + ReLU, `linear` Wz, `max` over the 32 slots (dim -2), `linear` Wy + b1 + ReLU, untilize. HiFi4 required (the identity blocks); `Precision.activations` = the dtype of the hidden `h` / `z` (the output rows are always bf16); `.release()`, `.describe()` | +| `features_shape(P)` -> `(1, 1, P, 512)`; `pack_features(features (P_kept, K, F), P, *, out=None)` | the device input: one 1 KB row per pillar (32 slots x 16 features, zero-padded), rows past `P_kept` zero; **K < 32 slots: slots K..31 repeat slot 0** (TransFusion's 20-point pillars: the slot maxima are unchanged, zero rows would add `relu(b0)`) | +| `pack_inputs(features, canvas_index, num_pillars, *, capacity)` | `{"features", "canvas_index"}` as numpy (tests and tools); the index range-checked (C19 `check_index`, sentinel = `capacity`) | +| `FrameStaging(num_pillars, num_cells)` | persistent bf16 `features` / uint32 `canvas_index` host tensors (`tensors.HostStaging`); `.write(features, canvas_index, num_pillars=None)` converts only the kept pillars (torch fp32 -> bf16 RNE, K < 32 slot rule), zeroes the rows the previous frame used beyond them, copies the checked index and returns the `TraceRunner` inputs; `.zero_copy` (False: falls back to `pack_features` + `HostStaging.write`, same values) | + +Device form (exact rewrites, also in floating point): one pillar = one 32-row tile, so a slot max is a tile-column +reduction; `relu(. + c)` is monotonic, so the per-pillar term `c = m @ W1b + b1` is added after the max (no slot +broadcast); `h` and the max ride through the matmuls in identity blocks (x * 1.0 is exact with HiFi4). Every pillar +row is computed, also past the frame's pillar count: the scatter index never points there, so no count reaches the +device. Precision: the bf16 rounding of the hidden `h` / `z` is the PFN's largest error term (PointPainting's CPU +emulation on its golden frames: output rel. L2 0.0045 with bf16 hidden tensors, 0.0020 with fp32 hidden tensors and +TF32 weights, 0.0017 for the bf16 output rounding alone), hence the `a=fp32` option; a model with large-magnitude input +columns (absolute x / y up to 121.6 m) can rewrite its first layer around the pillar centre on the host (PointPainting +`tt/pfn_input.py`, a promotion candidate) and feed this module unchanged. + +Device results (`tests/device/test_pillars_device.py`, TraceRunner: eager x2, strict capture, replay on input set B, +replay == eager bit for bit; the 512-pillar cases first under `TT_METAL_WATCHER=2`, clean; random weights, bf16 +inputs, vs the literal float64 encoder on the same inputs; `logs/pointpainting/m1/j1_pillars_*.log`, results +`logs/ttaw/pillars_device_results.json`): CenterPoint widths (9 -> 16, 32 -> 32) PCC 0.999996 / rel. L2 0.0025; +PointPainting (16 -> 16, 32 -> 32) bf16 hidden 0.999996 / 0.0025, fp32 hidden 0.9999975 / 0.0024; TransFusion (11 -> +32, 64 -> 64, 20-point pillars) bf16 hidden 0.999995 / 0.0027, fp32 hidden 0.999997 / 0.0025; the real capacity +(40,000 pillars, 28,137 kept, fp32 hidden) 0.9999975; `FrameStaging`'s zero-copy path equals `pack_features` + bf16 +RNE for a large frame, a smaller one (its stale rows zeroed) and 20-point pillars. The PointPainting bundle runs it +with its real weights (PFN PCC 0.99999 on two golden frames, `bundles/pointpainting-p150/PORT_LOG.md`). Host tests: +`tests/host/test_pillars_host.py` (12; fake ttnn: the plan exact against the literal encoder for the CP / PP / TF +widths and the 20-slot rule, packing, the module through `TraceRunner` replay == eager with bf16 and fp32 hidden +tensors, `FrameStaging`'s fallback path). + +--------------------------------------------------------------------------------------------------------------------- + +## 23. Segment reductions: K1 `segment_reduce` and the log-step fallback (`ttaw.ops.segment`) + +Built by the FRNet port (PLAN.md 1.3 row K1; owned by it) for FRNet's six `ScatterElements(max)` sites (points +sorted by frustum cell); PTv3 pooling and BEVPool (sum, optimization phase) are of the same form once the host sorts +the rows. A *segmented* tensor holds rows sorted by segment id: segment `s` is the contiguous row range +`[off[s], off[s + 1])` (CSR offsets, non-decreasing); rows past `off[-1]` are padding that nothing reads. + +```python +from .ttaw.ops import segment as seg + +off = seg.segment_offsets(seg_id_sorted, num_segments) # host CSR offsets, uint32 [S + 1] +row = seg.check_offsets(off, num_segments=S, num_rows=N_cap, length=S + 1) # [1, K] for the device +k1 = seg.SegmentReduce(seg.SegmentReduceSpec(num_rows=N_cap, channels=128, num_segments=M_cap + 1, + input_dtype="float32", input_layout="tile", + output_dtype="bfloat16")) +runner.add_input("offsets", init=row, dtype="uint32") # RT-dev: rewritten per frame +def forward(ctx): # a TraceRunner variant + table = k1(point_rows_fp32_tile, ctx["offsets"]) # [1, 1, M_cap + 1, 128] bf16 ROW_MAJOR; empty rows -> 0 + return gather_rows(ctx["pix2vox"], table, sentinel=M_cap) # C19: straight into a gather table +``` + +| name | meaning | +|---|---| +| `SegmentReduceSpec(num_rows, channels, num_segments, input_dtype="float32", input_layout="tile", output_dtype="float32", op="max", empty_value=0.0, workers_per_core=2)` | static shape of one call site (COMPILE). `num_rows % 32 == 0`, `channels % 32 == 0`; bf16 / fp32 in, TILE or ROW_MAJOR; bf16 (round to nearest even) / fp32 out; `op` `"max"` only on the device (`sum` / `mean` raise `NotImplementedError`: the data-movement RISC-Vs have no float unit, see below); `.l1_bytes_per_worker()`, `.check_l1(budget)` | +| `SegmentReduce(spec, *, name, grid=None)` | `__call__(x, offsets, *, output=None)` -> `[1, 1, num_segments, C]` ROW_MAJOR DRAM (allocated per call, or the persistent `output` of that spec): ONE `ttnn.generic_op`, eager or inside a capture. `x` `[1, 1, num_rows, C]` DRAM interleaved of the spec's dtype / layout; `offsets` uint32 ROW_MAJOR `[1, K >= num_segments + 1]` DRAM. `program(x, offsets, out)` (the `ProgramDescriptor`), `allocate_output(device)`, `output_spec()`, `describe()`. `grid` overrides the device grid (tests: one core) | +| `segment_reduce(x, offsets, *, num_segments, op="max", output_dtype=None, empty_value=0.0, output=None)` | one-shot form (spec from the tensors) | +| `logstep_segment_max(x, shift_tables, last_rows, *, layout="tile")` | the exact stock-op fallback (PLAN.md 1.3): `R` rounds `x = max(x, x[idx_r])` (untilize + `ttnn.embedding` + `ttnn.maximum`: 3 programs each), then a PADDED gather of `last_rows` from the result with a zero row appended at `R` (empty segments -> 0). bf16 in / out only (`ttnn.embedding` tables are bf16); exact for segments of up to `2^R` rows; `R` is COMPILE (bucket it per frame) | +| `segment_offsets(seg_id, num_segments, *, num_rows=None)`, `check_offsets(offsets, *, num_segments, num_rows, length=None)` | host CSR offsets from sorted ids (ids `>= num_segments` = trailing padding); validation (starts at 0, non-decreasing, ends `<= num_rows`) into the device `[1, length]` row | +| `segment_rounds(max_len)`, `segment_shift_tables(seg_id, rounds)`, `segment_last_rows(offsets, *, empty_row)` | the fallback's tables: `ceil(log2 L)`; `idx_r[i] = i - 2^r` in the same segment, else `i`; each segment's last row (`empty_row` = the row count: the appended zero row) | +| `segment_reduce_numpy(x, offsets, num_segments=None, *, op="max"\|"min"\|"sum"\|"mean", empty_value=0.0)`, `logstep_segment_max_numpy(x, shift_tables, last_rows)` | oracles (max / min exact in the input dtype; sum / mean in float64) | +| `worker_segments(offsets, num_segments, num_workers, *, num_rows=None)` | the kernel's work split, host twin (tests, diagnostics) | +| `OPS`, `DEVICE_OPS`, `KERNEL` | `("max", "min", "sum", "mean")`, `("max",)`, `"segment_reduce_dm.cpp"` | + +**Kernel** (`ops/kernels/segment_reduce_dm.cpp`, one source for both data-movement RISC-Vs of every core). Exact by +construction: the maximum is taken on integer keys of the float bit patterns (`key = bits ^ ((bits >> 31) & +0x7fffffff)`: signed order == float order, -0 < +0; NaN unsupported), so bf16 / fp32 results are bit-exact, and an +fp32 input with a bf16 output equals rounding first (RNE is monotonic). Work split without per-core runtime args +(probe P3 rule; common args = the three buffer addresses): worker `w` (= core index x 2 + RISC-V) owns the segments +whose cost `off[s] + s` (rows + one output row) lies in `[w T / W, (w + 1) T / W)`, found on the device by two binary +searches over the offsets (64-byte DRAM probes), so nothing in the host tables depends on the grid and ETH (12x10) +and WORKER (11x10) give identical outputs. Each worker streams its contiguous rows in units of 32 (one tile-row in +TILE mode; double-buffered NOC reads), reduces into an L1 accumulator and writes one output row per segment through +a ring of 8 staging rows; empty segments get `empty_value` (FRNet's frustum2pixel zero row comes out of the kernel). +Offsets are clamped to `[0, num_rows]` and to non-decreasing order, so a corrupt table cannot make it read outside +`x`. L1 per worker: 2 units + 1 KiB offsets window + 128 B probe + C x 4 B accumulator + 8 output rows (FRNet's +widest site, 256 fp32 channels: 75 KiB). Hang protocol (PLAN.md 4.4): no CB push / pop (CBs are plain scratch), no +multicast, no semaphores; DRAM reads keep the source's 64-byte alignment offset in L1 (Blackhole +`NOC_DRAM_READ_ALIGNMENT_BYTES`), every read is waited on (the barrier also invalidates the Blackhole L1 cache) +before use, staging rows are reused after `noc_async_writes_flushed`, and the kernel ends with both barriers. + +Why max only: the data-movement RISC-Vs (rv32im + Zba / Zbb, no F extension) compare integers natively (Zbb `max`) +but would emulate float adds in software; a `sum` / `mean` K1 needs the compute engine (the optimization-phase +consumers, PTv3 / BEVPool, add it with their device tests). + +Device results (`tests/device/test_segment_reduce_device.py`, the FRNet port's jobs; results +`logs/ttaw/k1_segment_reduce_device_results.json`): tiny bring-up under `TT_METAL_WATCHER=2` clean (64 x 32, 6 +segments incl. empty / 1-row / tile-row-crossing ones; one core with 1 and 2 workers; eager x2 and a 2-call trace +replayed on new data, bit-exact); FRNet's real sites through `TraceRunner` with the offsets rewritten per frame (10 +replays x 5 rewrites), every output bit-exact vs the oracle and replay == eager, 240 workers (12x10 ETH grid, 2 per +core): OT128 encoder site 160,000 x 256 fp32 -> 60,000 fp32 rows 5.52 ms per replay, OT128 point sites 160,000 x 128 +fp32 -> 60,001 bf16 rows 2.89 ms, QT128-like dense segments (6,255 segments, one of 500 rows) 2.22 ms, 2,048-row cases +0.04-0.08 ms; the log-step fallback equals K1 bit for bit (4 rounds on 4,096 rows; 7 rounds on 120,000 rows with +111-row segments) (`logs/frnet/m1a/j1_bringup_and_graph.log`, watcher log `logs/frnet/m1a/j1_tiny_watcher.log`, +`logs/ttaw/k1_segment_reduce_device_results.json`). Inside the FRNet graph (six sites per frame, OT128 sample): every +tap downstream PCC >= 0.9995 vs the fp32 reference (`bundles/frnet-p150/PORT_LOG.md`). A 2,000-launch soak +(`test_soak`) runs in the port's next job. Host tests: +`tests/host/test_segment_host.py` with `tests/host/fake_ttnn_segment.py` (program-descriptor records, a +`generic_op` that emulates the kernel from its compile-time args and checks the P3 rules and the split, `maximum`, +`hardswish`, `softmax`, `addcmul`). diff --git a/code/tt_diffusion_planner/ttaw/VENDORED.json b/code/tt_diffusion_planner/ttaw/VENDORED.json new file mode 100644 index 0000000000000000000000000000000000000000..3b69756fe708e05d47d45c5cf5d17dfa381a7d1a --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/VENDORED.json @@ -0,0 +1,55 @@ +{ + "package": "ttaw", + "version": "0.20.0", + "source": "common/ttaw", + "source_commit": "89dec49a7e66836ee745660e1cfedf228a4c070b", + "source_ref": "HEAD", + "source_tree": "commit", + "source_dirty": false, + "vendored_at": "2026-10-09T04:21:08+00:00", + "files": { + "API.md": "643842c2199e6079269f7ec5f3ef0fa9eeebcaab8289b9c4a6b1c3a3b090d9e2", + "__init__.py": "fb754d83e40d05651a667c469e1c42ad1d5d0b3f779bf9710fae16c88b1bbcfe", + "api_base.py": "189ee4215f937b812afee40d076bb948861a489e0b9e42a102bba311e0536179", + "decode.py": "f8aa2efa7ec664da7b8af94bed2d5771552b08b1a6d0bf97cbafcd18ea805c90", + "device.py": "06357ba55c9f7a4254982fb1effc15b9cb85d3158c33afc6d6c5447071d918ce", + "geometry.py": "087dcf7f43ef187fb051a504fe15365985c2c543eb3e05554a9017bcef90fcd5", + "golden.py": "d675431f052bfd01438aa79d3a5e281380a498b1f762e70b0ae0f3b4f194c7ad", + "image.py": "0a755647744d15f84077fe396d2514ba996a3ca056d172327874df2f21b52d76", + "image_area.py": "8d9e171167cf3fe2b59505ae9e5661cf1c032e21f282aa3e78361b4140e3d98c", + "image_linear.py": "debd3ade236a69f74859ae91ac9013c4757198ac29508f06f85d195c38e49d31", + "image_triangle.py": "b3f2f08a3c2512c4df28b987d8002358ef5b7556bc1ad29f5b91c956b7d154c2", + "io.py": "c2ac25cb589351a47b26688759d198204b5f0c25fd92f2a5c1c6428cea5a149e", + "knobs.py": "888783782c8f2e230538f416c360791f5cf8b21e5330414596fda7898308a279", + "metrics.py": "6d3a6b36d287aba05537b57d3c7560fd93537063e6b5c592492ed7d6e410e9e0", + "models/__init__.py": "2d2783c72c0382a0ffa34f71447fe75279fd357099b663c506af76d6bc2c2ec5", + "models/centerhead.py": "0f404b39b5f063e8d29e546887e03c3681630d1f4d98fffab693c501fc92cff6", + "models/pillars.py": "a7357e3c43365608fe46bbb6a3cd7cef86b2afaa751a28c0dfa93b617334dc1c", + "models/resnet.py": "51183499e25f8146c68a417d2ac1a176c8f941ca224b58609c34bb356fe500b0", + "models/second.py": "8e98236f9903c8ce52b79ccc18869b5a8cccae8401baa9f2294c4dd2d7b2111d", + "models/transfusion_head.py": "e040f86979278aa61b420a6d077a854482967e73cfa419926f6df721eb3c812a", + "nms.py": "c3c9e63b201c68211fcb6a9a2844e63f7233212354e9b61eb869d9ae12097116", + "ops/__init__.py": "6d6aa15abd7f3cded433e083988a6f420f351e2bf0d7765d9e8c9c05d26fec64", + "ops/attention.py": "73e170cbd483590de08867881a282816b28d359bde91ff70de93dec6c786e404", + "ops/conv.py": "d8c5a31ef710aeb9568e638b721d4255f63743a1b88f822e6f411dee4e158532", + "ops/deform.py": "ac10088af29bc83fd66fab02712c836ffa5795ec4f7304846722ca54181897cd", + "ops/gather.py": "e2934f898b49088c387e8e89db31681257cdba03204efbee4160e055a38f0280", + "ops/heatmap.py": "05f13986f0ed7c59e1d80877ab5ca8cca7528e30c60e43dc489086f58732f596", + "ops/kernels/segment_reduce_dm.cpp": "a313c69865990decdccebcbd0ef91253b0a0f1b4f97370212bdbd89a11b0305f", + "ops/segment.py": "db4ea1a73e24868956240474a473f95ac94bd00dd87ec8ba6c2fb5cbce828457", + "ops/topk.py": "a87817a35afb3448eda140376d8170d69ddc9acf74b3717c62a2987fc4bcddf4", + "ops/upsample.py": "bbea4ef32298277393c73d427fa00aa2fd4d68d478c770de731010ab4c9e34dd", + "outputs.py": "e345695de6a888ee617afbb1a79bfd4251e768a31446fadf7d2ca6aba330d6ca", + "pointcloud.py": "fba6368408de16d258946b2814e7e1e286495c870750dafe52c92f7eec985758", + "precision.py": "8a01f1d048c3d317cc3caa095421e9695c2f6b7fdad2623a8d0c26364bdeaf3d", + "profiling.py": "c65aefbf0c2df111246f927f20220619686c51ad3b959ddb030655e97d8461f4", + "server/__init__.py": "3bbfe7beb80bf1b7a983c264e1ae3f08e162da17289b397ff0c4f425bb4e689a", + "server/app.py": "23cbf230d7393f743527fda1706f180c90fb8c6c2089fbf6559bd29d10d951dc", + "server/client.py": "e78f07ebdb132cb06c069259903dc7f971044f75f4918aacd184e5a59d7ca987", + "server/smoke.py": "acc547a03bcfe719ca1b5fba4e551d0dadecf77505f8e39848747b32c082b33b", + "tensors.py": "8abc00a10f5d0524134b49225a6ee4aba1104bfc19fc875eb777c3db73da5fbc", + "trace.py": "56215f3cdcad85c4d7063ff6e4a6a6acc5d4f90c2d855b2524898aaefd4d0ad6", + "voxelize.py": "88183715a10be92d489d35acf88defa4eaf47763fcb0356a71d64d97acc859b6", + "weights.py": "c8857817b7273e48278bb8194cece7fc38d99722119caacf8eb8fb8689648fad" + } +} diff --git a/code/tt_diffusion_planner/ttaw/__init__.py b/code/tt_diffusion_planner/ttaw/__init__.py new file mode 100644 index 0000000000000000000000000000000000000000..c47c91d101e964a2a347823ed192d71a7255a9ef --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/__init__.py @@ -0,0 +1,103 @@ +# SPDX-License-Identifier: Apache-2.0 +"""ttaw: shared infrastructure of the Autoware ports to one Tenstorrent Blackhole p150 (ttnn / tt-metalium). + +The package is the source of truth under ``common/ttaw`` and is vendored into every bundle as the sub-package +``code//ttaw`` by ``common/tools/vendor.py``. Rules that keep one tree valid in both places: + +- relative imports only (``from .device import open_device``), never ``import ttaw``; +- importing any module has no side effects: no ``import ttnn`` / ``torch`` at module level, no device, no network, + no environment reads (ttnn and torch are imported inside the functions that need them); +- optimization knobs are read once at model build (:mod:`.knobs`) and each has an env A/B switch; +- custom kernel ``.cpp`` files are package data, located with :func:`.ops.kernel_path`. + +Modules (see ``API.md`` for the full guide): + +================== ====== ============================================================================= +module id what it provides +================== ====== ============================================================================= +``device`` C01 ``open_device`` (ETH dispatch default, ``auto`` patch detection), ``describe_device`` +``trace`` C02 ``TraceRunner``: persistent I/O, warm-up, capture, variants, 1CQ / 2CQ, state +``tensors`` C03 tile padding, layout / dtype helpers, fp32 islands, zero-copy host staging buffers +``weights`` C04 ``OnnxWeights``, BN folding in fp64, safetensors / ``.pth`` loading, weight cache +``precision`` C05 explicit ``compute_kernel_config`` factory and per-module policy tables +``knobs`` -- env A/B knobs read once at build +``metrics`` C06 PCC variants, argmax agreement, IoU, top-K overlap, detection matching, ADE/FDE +``golden`` C06 golden ``.npz`` files, tap registry, gate registry that never loosens, reports +``profiling`` C07 stage bench (p50/p99), Tracy signposts, ops-CSV summarizer, AICLK sampler +``io`` C08 point cloud / image / calibration decoding, output encoders +``outputs`` C08 ``Detections3D`` ``Detections2D`` ``Segmentation3D`` ``Mask2D`` ``Trajectory`` +``api_base`` C08 ``ModelBase`` (from_pretrained, weights resolution, warm-up, lock, info, close) +``server`` C08 FastAPI app factory (``server.app``); stdlib client, smoke and reference checks + (``server.client``, ``server.smoke``) +``image`` C14 bit-exact Autoware image pre-processing from cached LUTs (YOLOX letterbox, BEVDet crop) +``image_area`` C14 METEOR preset: OpenCV ``INTER_AREA`` down-scaling of uint8 frames, bit-exact +``image_linear`` C14 SceneSeg preset: OpenCV ``INTER_LINEAR`` / ``INTER_NEAREST`` resizes of uint8 frames and the + VisionPilot normalisation (BGR / RGB), bit-exact +``image_triangle`` C14 StreamPETR preset: Autoware's anti-aliased triangle resize + crop + normalise kernel, + bit-exact (normalisation presets autoware_main / autoware_0.52 / awml_training) +``geometry`` C10 rigid transforms (float64), Autoware's float32 sweep transform, yaw conventions +``pointcloud`` C11 Autoware multi-sweep densification state machine (per stream), hygiene masks +``voxelize`` C12 deterministic pillars, 9/10/11-feature decoration, canvas scatter / gather index +``nms`` C15 circle NMS, perception_utils IoU-BEV NMS, area class remapper +``decode`` C16 CenterHead dense decode (and at given cells), Autoware DetectedObject mapping +``ops`` -- kernel package-data helpers (custom ``generic_op`` kernels live in ``ops/kernels``) +``ops.conv`` C17 conv2d builders (weights prepared once, fused activations, DRAM slicing), K-split, + ConvTranspose k == s, feature-map glue +``ops.upsample`` C18 nearest up-sampling, ``F.interpolate`` matrices, separable device resize +``ops.attention`` C20 SDPA wrapper (non-causal, explicit scale, chunk table, aligned masked keys), heads +``ops.gather`` C19 gather-form scatter (zero sentinel row + TILE / PADDED ``ttnn.embedding``), row gathers +``ops.topk`` C21 top-k proposal selection (class-major ``topk_large_indices`` row, exact decode, set metrics) +``ops.heatmap`` C22 sigmoid + 3x3 local max of TF / BF / PT heatmaps (one exact masked chain) +``ops.deform`` C23 grid_sample helpers: fp32 affine / resize grids (exact identity), P9 rules, kernel emulation +``ops.segment`` K1 per-segment max over rows sorted by segment (``generic_op`` kernel, exact, no + per-core args), the log-step stock-op fallback, CSR / shift tables, oracles +``models`` C24 C25 device sub-networks: ``models.second`` (SECOND + SECONDFPN), ``models.pillars`` + C27 C26 (PillarFeatureNet + its input staging), ``models.centerhead`` (CenterPoint dense + head: K-split shared conv, merged heads), ``models.resnet`` (ResNet + bottleneck / basic-block builders, cameras as batch), ``models.transfusion_head`` + (TransFusion query head: selection, query init, decoder layer, merged heads) +================== ====== ============================================================================= +""" + +__version__ = "0.20.0" + +__all__ = [ + "__version__", + "api_base", + "decode", + "device", + "geometry", + "golden", + "image", + "image_area", + "image_linear", + "image_triangle", + "io", + "knobs", + "metrics", + "models", + "nms", + "ops", + "outputs", + "pointcloud", + "precision", + "profiling", + "server", + "tensors", + "trace", + "voxelize", + "weights", +] + + +def __getattr__(name: str): + """Load submodules on first attribute access (``ttaw.trace``), keeping ``import ttaw`` free of side effects.""" + if name in __all__ and name != "__version__": + import importlib + + return importlib.import_module(f".{name}", __name__) + raise AttributeError(f"module {__name__!r} has no attribute {name!r}") + + +def __dir__(): + return sorted(set(globals()) | set(__all__)) diff --git a/code/tt_diffusion_planner/ttaw/api_base.py b/code/tt_diffusion_planner/ttaw/api_base.py new file mode 100644 index 0000000000000000000000000000000000000000..abee4fc03eec88d852e5b832f2c6b6ae8d41715b --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/api_base.py @@ -0,0 +1,383 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C08 ``ModelBase``: the Python API contract shared by every bundle (BUNDLE_CONVENTIONS.md section 8). + +A bundle's ``api.py`` subclasses :class:`ModelBase`, sets the identity class attributes and implements six hooks:: + + class CenterPoint(ModelBase): + MODEL_NAME = "centerpoint-p150" + ENV_PREFIX = "CENTERPOINT" + DEFAULT_REPO = "AutowareFoundation/lidar_centerpoint" + DEFAULT_TAG, DEFAULT_REVISION = "v4.1", "494c8171def40bd36cc2feb323e0a5acbfab132b" + ALLOW_PATTERNS = ["base/*", "tiny/*"] + VARIANTS, DEFAULT_VARIANT = ("base", "tiny"), "base" + LABELS = ("CAR", "TRUCK", "BUS", "BICYCLE", "PEDESTRIAN") + RUNTIME_PARAMS = {"score_threshold": (float, 0.0, 1.0, 0.35)} + DEVICE_DEFAULTS = {"num_command_queues": 1, "trace_region_size": 64 << 20} + + def _build(self): ... # weights -> device tensors, TraceRunner + variants (no capture) + def _warm_one(self, v): ... # capture the variant's traces (TraceRunner.capture) + def _prepare(self, **kw): ... # host pre-processing -> host buffers + def _forward(self, prep): ... # upload + replay + read (under the model lock) + def _postprocess(self, raw, prep, params): ... # -> an ttaw.outputs object + def _release(self): ... # TraceRunner.release() + +Contract: ``from_pretrained`` resolves the pinned weights *before* opening the chip (a Hub problem must not claim +the device), opens it (ETH dispatch, 12x10), builds, and warms every default variant, so the first call is as fast +as the next. Calls are serialized by a re-entrant lock; ``close()`` is idempotent and also runs at interpreter exit. +Importing this module has no side effects (no ttnn, no torch, no network). +""" +from __future__ import annotations + +import atexit +import logging +import math +import os +import threading +import time +import weakref +from pathlib import Path +from typing import Any, ClassVar, Dict, List, Mapping, Optional, Sequence, Tuple + +import numpy as np + +from . import io as tio +from .device import DeviceConfig, close_device, describe_device + +__all__ = ["ModelBase", "resolve_weights"] + +log = logging.getLogger(__name__) + +# the input keyword arguments of ModelBase.__call__ (never runtime params) +_CALL_INPUTS = ("points", "sweeps", "images", "calibration", "stream", "inputs") + + +def resolve_weights(model_id: str, *, revision: Optional[str] = None, allow_patterns: Optional[Sequence[str]] = None, + weights_dir: Optional[str] = None, env_var: Optional[str] = None) -> Path: + """Local directory holding the weights files. + + Order: ``weights_dir`` > ``$`` > ``model_id`` if it is a local directory > the Hugging Face snapshot of + ``model_id`` at ``revision`` restricted to ``allow_patterns``. If the Hub is unreachable the cached snapshot is + used (``local_files_only=True``). Never needs a token for public repos; never uploads.""" + for cand in (weights_dir, os.environ.get(env_var) if env_var else None): + if cand: + p = Path(cand).expanduser() + if not p.is_dir(): + raise FileNotFoundError(f"weights directory {p} does not exist") + return p + if Path(model_id).expanduser().is_dir(): + return Path(model_id).expanduser() + from huggingface_hub import snapshot_download + + patterns = list(allow_patterns) if allow_patterns else None + try: + return Path(snapshot_download(model_id, revision=revision, allow_patterns=patterns)) + except Exception as e: # noqa: BLE001 -- offline or rate limited: fall back to the local cache + log.warning("Hub not reachable for %s@%s (%s: %s); trying the local HF cache", model_id, revision, + type(e).__name__, e) + return Path(snapshot_download(model_id, revision=revision, allow_patterns=patterns, local_files_only=True)) + + +_LIVE: "weakref.WeakSet" = weakref.WeakSet() + + +@atexit.register +def _close_all() -> None: + """A model that was not closed is closed when Python exits, so the chip is left clean.""" + for model in list(_LIVE): + try: + model.close() + except Exception: # noqa: BLE001 -- best effort at exit + pass + + +class ModelBase: + """Base class of every bundle's model class. Create instances with :meth:`from_pretrained`.""" + + # ---- identity (override in the bundle) --------------------------------------------------------------------- + MODEL_NAME: ClassVar[str] = "model" + ENV_PREFIX: ClassVar[str] = "MODEL" # _WEIGHTS_DIR, _DISPATCH, _NUM_CQS, ... + DEFAULT_REPO: ClassVar[str] = "" + DEFAULT_TAG: ClassVar[Optional[str]] = None + DEFAULT_REVISION: ClassVar[Optional[str]] = None # the commit DEFAULT_TAG points to (tags can move) + ALLOW_PATTERNS: ClassVar[Optional[Sequence[str]]] = None + WEIGHTS_LICENSE: ClassVar[str] = "Apache-2.0 (AutowareFoundation model card)" + VARIANTS: ClassVar[Sequence[str]] = ("default",) + DEFAULT_VARIANT: ClassVar[str] = "default" + INPUT_KIND: ClassVar[str] = "lidar" # lidar | camera | multicam | lidar+multicam | planner + CAMERA_ORDER: ClassVar[Sequence[str]] = () + REQUIRE_CALIBRATION: ClassVar[Optional[bool]] = None # None: True for multicam / lidar+multicam + POINT_FIELDS: ClassVar[Sequence[str]] = tio.DEFAULT_POINT_FIELDS + LABELS: ClassVar[Sequence[str]] = () + # name -> (type, min, max, default): host-side knobs only (BUNDLE_CONVENTIONS.md section 9); min / max may be + # None for an open side + RUNTIME_PARAMS: ClassVar[Mapping[str, Tuple[type, Any, Any, Any]]] = {} + # model-specific input keyword arguments passed to _prepare (not runtime params), e.g. PointPainting ("rois",): + # the server's ServerSpec.decode_extra adds them to the call kwargs + EXTRA_INPUTS: ClassVar[Sequence[str]] = () + # planner-style named inputs: name -> (shape with None for free dims, dtype); when set, ``inputs=`` is decoded + # and checked with io.load_named_arrays on every call (API and server alike) + INPUT_SCHEMA: ClassVar[Optional[Mapping[str, Tuple[Sequence[Optional[int]], Any]]]] = None + # defaults for ttaw.device.DeviceConfig; _* environment variables override them + DEVICE_DEFAULTS: ClassVar[Mapping[str, Any]] = {} + + def __init__(self, *_args: Any, **_kwargs: Any): + raise TypeError(f"use {type(self).__name__}.from_pretrained(...)") + + # ---- weights / device ------------------------------------------------------------------------------------ + @classmethod + def requires_calibration(cls) -> bool: + if cls.REQUIRE_CALIBRATION is not None: + return bool(cls.REQUIRE_CALIBRATION) + return cls.INPUT_KIND in ("multicam", "lidar+multicam") + + @classmethod + def resolve_weights(cls, model_id: Optional[str] = None, revision: Optional[str] = None, + weights_dir: Optional[str] = None) -> Path: + """:func:`resolve_weights` with this model's defaults (pinned revision for the default repo).""" + model_id = model_id or cls.DEFAULT_REPO + rev = revision or (cls.DEFAULT_REVISION if model_id == cls.DEFAULT_REPO else None) + return resolve_weights(model_id, revision=rev, allow_patterns=cls.ALLOW_PATTERNS, weights_dir=weights_dir, + env_var=f"{cls.ENV_PREFIX}_WEIGHTS_DIR") + + @classmethod + def device_config(cls, **overrides: Any) -> DeviceConfig: + """``DEVICE_DEFAULTS`` < ``_*`` / ``TT_DEVICE_ID`` environment < explicit non-None ``overrides`` + (``device_id``, ``dispatch``, ``num_command_queues``, ...).""" + base = DeviceConfig.from_env(cls.ENV_PREFIX, **dict(cls.DEVICE_DEFAULTS)) + values = {k: getattr(base, k) for k in base.__dataclass_fields__} + values.update({k: v for k, v in overrides.items() if v is not None}) + return DeviceConfig(**values) + + @classmethod + def from_pretrained(cls, model_id: Optional[str] = None, *, revision: Optional[str] = None, + variant: Optional[str] = None, device_id: Optional[int] = None, device: Any = None, + dispatch: Optional[str] = None, num_command_queues: Optional[int] = None, + weights_dir: Optional[str] = None, warmup_variants: Any = "default", verbose: bool = False, + **compile_params: Any) -> "ModelBase": + """Resolve weights, open the chip (unless ``device`` is given; ``close()`` then leaves it open), build the + graph and capture the traces of ``warmup_variants`` (``"default"``, ``"none"`` or a list). + + ``variant``: one of ``VARIANTS`` (load-time; default ``_VARIANT`` or ``DEFAULT_VARIANT``). + ``device_id``: default ``TT_DEVICE_ID`` or 0. ``dispatch``: ``"eth"`` / ``"worker"`` / ``"auto"`` (default + ``_DISPATCH`` or eth). ``num_command_queues``: default ``_NUM_CQS`` or ``DEVICE_DEFAULTS``. + ``compile_params``: shape-defining options, validated in ``_build``.""" + variant = variant or os.environ.get(f"{cls.ENV_PREFIX}_VARIANT") or cls.DEFAULT_VARIANT + if variant not in cls.VARIANTS: + raise ValueError(f"variant={variant!r}: expected one of {list(cls.VARIANTS)}") + clash = sorted(set(cls.RUNTIME_PARAMS) & (set(_CALL_INPUTS) | set(cls.EXTRA_INPUTS))) + if clash: + raise TypeError(f"{cls.__name__}: RUNTIME_PARAMS {clash} clash with input keyword arguments") + self = cls.__new__(cls) + self._lock = threading.RLock() + self._closed = False + self._owns_device = device is None + self.device = None + self.variant = variant + self.compile_params = dict(compile_params) + self.verbose = verbose + self.warmup_ms: Dict[str, float] = {} + self.warm_variants: List[Any] = [] + self.device_info: Dict[str, Any] = {} + model_id = model_id or cls.DEFAULT_REPO + rev = revision or (cls.DEFAULT_REVISION if model_id == cls.DEFAULT_REPO else None) + t0 = time.perf_counter() + log.info("Loading weights %s@%s (variant %s)", model_id, rev, variant) + self.weights_path = cls.resolve_weights(model_id, revision, weights_dir) + self.weights = {"repo": model_id, "tag": cls.DEFAULT_TAG if model_id == cls.DEFAULT_REPO else None, + "revision": rev, "path": str(self.weights_path), "license": cls.WEIGHTS_LICENSE} + self.warmup_ms["weights"] = (time.perf_counter() - t0) * 1e3 + try: + if device is None: + cfg = cls.device_config(device_id=device_id, dispatch=dispatch, num_command_queues=num_command_queues) + log.info("Opening device %d (dispatch %s, %d CQ)", cfg.device_id, cfg.dispatch, cfg.num_command_queues) + device = cfg.open() + device_id = cfg.device_id + self.device = device + self.device_info = describe_device(device, device_id) + t1 = time.perf_counter() + self._build() + self.warmup_ms["build"] = (time.perf_counter() - t1) * 1e3 + self.warmup(warmup_variants) + except BaseException: + self.close() + raise + self.warmup_ms["total"] = (time.perf_counter() - t0) * 1e3 + _LIVE.add(self) + return self + + # ---- warm-up --------------------------------------------------------------------------------------------- + def default_warmup_variants(self) -> List[Any]: + """The trace variants captured by default (e.g. pillar-count buckets). Override per port.""" + return [{"variant": self.variant}] + + def warmup(self, variants: Any = "default") -> Dict[str, Any]: + """Compile + capture each variant (idempotent). Returns ``{"ms": ..., "variants": [...]}``.""" + self._check_open() + if variants == "default": + todo = self.default_warmup_variants() + elif variants in (None, "none"): + todo = [] + else: + todo = list(variants) + t0 = time.perf_counter() + with self._lock: + for v in todo: + if v in self.warm_variants: + continue + log.info("Warming up: capturing trace for %s", v) + self._warm_one(v) + self.warm_variants.append(v) + ms = (time.perf_counter() - t0) * 1e3 + self.warmup_ms["warmup"] = self.warmup_ms.get("warmup", 0.0) + ms + log.info("Warmup complete: %d variant(s) in %.0f ms", len(self.warm_variants), ms) + return {"ms": ms, "variants": list(self.warm_variants)} + + # ---- inference ------------------------------------------------------------------------------------------- + @classmethod + def validate_params(cls, params: Optional[Mapping[str, Any]]) -> Dict[str, Any]: + """Per-request knobs -> typed values with defaults. Unknown names, values outside ``[min, max]`` (either + bound may be None), non-booleans for ``bool``, and booleans or non-integral numbers for ``int`` raise + ``InputError``; numeric strings are accepted for numbers. ``None`` is accepted for a param whose default + is ``None`` (it means "not set"), so ``validate_params(validate_params(p)) == validate_params(p)``.""" + out = {k: spec[3] for k, spec in cls.RUNTIME_PARAMS.items()} + for k, v in (params or {}).items(): + if k not in cls.RUNTIME_PARAMS: + raise tio.InputError(f"unknown parameter {k!r}; allowed: {sorted(cls.RUNTIME_PARAMS)}") + typ, lo, hi, default = cls.RUNTIME_PARAMS[k] + if v is None and default is None: # "not set" (JSON null) = the default, so validation is idempotent: + out[k] = None # the server validates, then model(**params) validates again + continue + if typ is bool and not isinstance(v, bool): + raise tio.InputError(f"parameter {k!r} must be a boolean") + if typ in (int, float) and isinstance(v, bool): + raise tio.InputError(f"parameter {k!r} must be {typ.__name__}, not a boolean") + try: + if typ is int and not isinstance(v, int): + number = float(v) + if not number.is_integer(): + raise ValueError(v) + v = int(number) + else: + v = typ(v) + except (TypeError, ValueError, OverflowError): + raise tio.InputError(f"parameter {k!r} must be {typ.__name__}") from None + if typ is float and not math.isfinite(v): + raise tio.InputError(f"parameter {k!r} must be finite") + if (lo is not None and v < lo) or (hi is not None and v > hi): + raise tio.InputError(f"parameter {k!r}={v} is outside [{lo}, {hi}]") + out[k] = v + return out + + def __call__(self, points: Any = None, *, sweeps: Optional[Sequence] = None, images: Any = None, + calibration: Optional[Mapping] = None, stream: Optional[Mapping] = None, + inputs: Any = None, **params: Any): + """One frame (batch 1). ``points``: path / bytes / (N, C) array / PointCloud / JSON envelope. + ``sweeps``: ``[{"points", "time_lag_s", "T_current_from_sweep"}]``. ``images``: list of camera dicts or + CameraImage. ``calibration``: dict (or ``{"preset": name}``, resolved by the server). ``stream``: + ``{"id", "reset", "timestamp_s", "T_world_from_ego"}``. ``inputs``: named arrays (planner; checked + against ``INPUT_SCHEMA`` when the class sets one). Keyword arguments named in ``EXTRA_INPUTS`` go to + ``_prepare``; the others are ``RUNTIME_PARAMS``. Returns the model's output object with ``timing_ms``.""" + self._check_open() + extra = {k: params.pop(k) for k in list(params) if k in self.EXTRA_INPUTS} + p = self.validate_params(params) + if inputs is not None and self.INPUT_SCHEMA is not None: + inputs = tio.load_named_arrays(inputs, self.INPUT_SCHEMA) + with self._lock: + self._check_open() # close() may have run while this call waited for the lock + t0 = time.perf_counter() + prepared = self._prepare(points=points, sweeps=sweeps, images=images, calibration=calibration, + stream=stream, inputs=inputs, **extra) + t1 = time.perf_counter() + raw = self._forward(prepared) + t2 = time.perf_counter() + out = self._postprocess(raw, prepared, p) + t3 = time.perf_counter() + out.timing_ms.update({"preprocess": (t1 - t0) * 1e3, "device": (t2 - t1) * 1e3, + "postprocess": (t3 - t2) * 1e3, "total": (t3 - t0) * 1e3}) + return out + + predict = __call__ + + # ---- lifetime -------------------------------------------------------------------------------------------- + def extra_info(self) -> Dict[str, Any]: + """Port-specific additions to :attr:`info` (knobs, precision policy, TraceRunner.describe() ...).""" + return {} + + @property + def info(self) -> Dict[str, Any]: + cls = type(self) + info = {"model": cls.MODEL_NAME, "variant": getattr(self, "variant", None), + "weights": getattr(self, "weights", None), "device": getattr(self, "device_info", None), + "warm_variants": getattr(self, "warm_variants", []), "warmup_ms": getattr(self, "warmup_ms", {}), + "compile_params": getattr(self, "compile_params", {}), "input_kind": cls.INPUT_KIND, + "point_fields": list(cls.POINT_FIELDS), "camera_order": list(cls.CAMERA_ORDER), + "labels": list(cls.LABELS), "runtime_params": {k: v[3] for k, v in cls.RUNTIME_PARAMS.items()}, + "extra_inputs": list(cls.EXTRA_INPUTS)} + if cls.INPUT_SCHEMA is not None: + info["input_schema"] = {name: {"shape": [None if d is None else int(d) for d in shape], + "dtype": np.dtype(dtype).name} + for name, (shape, dtype) in cls.INPUT_SCHEMA.items()} + if not getattr(self, "_closed", True): + info.update(self.extra_info()) + return info + + @property + def closed(self) -> bool: + return getattr(self, "_closed", True) + + def _check_open(self) -> None: + if self.closed: + raise RuntimeError(f"{type(self).__name__} is closed") + + def close(self) -> None: + """Release traces and device tensors; close the chip if this model opened it. Idempotent.""" + if getattr(self, "_closed", True): + return + self._closed = True + _LIVE.discard(self) + with self._lock: + try: + self._release() + finally: + dev = getattr(self, "device", None) + if dev is not None and self._owns_device: + close_device(dev) + self.device = None + + def __enter__(self) -> "ModelBase": + return self + + def __exit__(self, *exc: Any) -> None: + self.close() + + def __repr__(self) -> str: + state = ("closed" if self.closed + else f"{self.device_info.get('dispatch')} {self.device_info.get('grid')}") + return f"<{type(self).__name__} {self.MODEL_NAME} variant={getattr(self, 'variant', '?')} {state}>" + + # ---- port-specific hooks --------------------------------------------------------------------------------- + def _build(self) -> None: + """Load weights from ``self.weights_path``, convert them to device tensors, build the ttnn graph and its + ``TraceRunner`` (inputs, params, states, variants). No capture here.""" + raise NotImplementedError("port-specific: build the ttnn graph") + + def _warm_one(self, variant: Any) -> None: + """Warm up and capture the traces of one variant (``TraceRunner.capture``).""" + raise NotImplementedError("port-specific: compile + capture one variant") + + def _prepare(self, **kwargs: Any) -> Any: + """Host pre-processing ported from the Autoware node -> host buffers. Raise ``InputError`` for client + mistakes (missing input, wrong shape).""" + raise NotImplementedError("port-specific: host pre-processing") + + def _forward(self, prepared: Any) -> Any: + """Upload into the persistent device inputs, replay the matching trace, read the outputs.""" + raise NotImplementedError("port-specific: traced device forward") + + def _postprocess(self, raw: Any, prepared: Any, params: Dict[str, Any]) -> Any: + """Host post-processing ported from Autoware -> an output object of :mod:`ttaw.outputs`.""" + raise NotImplementedError("port-specific: host post-processing") + + def _release(self) -> None: + """Release traces and persistent device tensors (``TraceRunner.release()``). Also called when + ``from_pretrained`` fails half-way, so tolerate a partially built model (``getattr(self, "runner", None)``).""" diff --git a/code/tt_diffusion_planner/ttaw/decode.py b/code/tt_diffusion_planner/ttaw/decode.py new file mode 100644 index 0000000000000000000000000000000000000000..6b11c799629da67c74eb2bd6698814a84a5315df --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/decode.py @@ -0,0 +1,424 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C16 decode: the dense CenterHead decoder of Autoware's lidar_centerpoint (also PointPainting's +image_projection_based_fusion), and the mapping of decoded boxes to Autoware ``DetectedObject`` fields. + +``decode_centerhead`` is ``generateBoxes3D_kernel`` + the ``is_score_keep`` compaction +(``autoware_lidar_centerpoint/lib/postprocess/postprocess_kernel.cu:52-178,220-226``), evaluated per BEV cell in +float32 as the kernel does (S:centerpoint:281-287): + +- ``score_c = sigmoid(heatmap[c])``; ``label`` = first class with the largest score (strict ``>`` scan); +- ``x = (voxel_x * ds) * (xi + reg[0]) + min_x``, ``y`` likewise with ``yi``: **no +0.5 cell offset**; +- the distance bin is the first ``i`` with ``sqrt(x^2 + y^2) < upper[i]``; beyond the last bin the cell is dropped; +- the cell is dropped when ``score < thresholds[bin][label]`` or ``sqrt(rot0^2 + rot1^2) < yaw_norm[label]``, and a + zero score is never kept; +- box: ``z = height[0]`` (box centre), ``width = exp(dim[0])``, ``length = exp(dim[1])``, ``height = exp(dim[2])`` + (the deployed ONNX regresses (w, l, h)), ``yaw = atan2(rot[0], rot[1])``, ``vel = vel[0:2]``. + +No max-pool peak finding and no top-K: every cell over threshold is a candidate, as in Autoware (S:centerpoint:310). +Thresholds outside [0, 1) are coerced to 0 (``centerpoint_config.hpp:79-93``): :meth:`CenterHeadDecodeConfig.create` +applies that. The decoder works on head maps from any source (ONNX Runtime, the torch reference, a device readback). + +Order: rows come out in cell order (``yi * W + xi``); :func:`sort_by_score` is the ``thrust::sort`` by descending +score, made deterministic (stable: ties keep the ascending cell order; thrust's sort is not stable). + +``decode_cells`` evaluates the same decoder at given cells and drops none: it keeps each cell's gate inputs +(distance bin, score threshold, yaw norm) and outcomes, so two head-map sources (device and fp32 reference) can be +compared at the same cells and every cell one keeps and the other drops can be explained. ``head_values_at`` stores +the maps at a set of cells (compact agreement goldens). + +``to_detected_objects`` is ``box3DToDetectedObject`` (``ros_utils.cpp:29-87``): the class name maps to an +``autoware_perception_msgs/ObjectClassification`` label (``getSemanticType``), ``yaw_ros = -yaw - pi/2`` (float), +dimensions (length, width, height), ``orientation_availability`` SIGN_UNKNOWN for car-like labels, and the twist in +the object frame when ``has_twist``. + +numpy only; no side effects on import. +""" +from __future__ import annotations + +import math +from dataclasses import dataclass, field, fields, replace +from typing import Any, Dict, List, Mapping, Optional, Sequence + +import numpy as np + +__all__ = [ + "AUTOWARE_LABELS", + "LABEL_IDS", + "SIGN_UNKNOWN", + "UNAVAILABLE", + "semantic_label", + "is_car_like", + "CenterHeadDecodeConfig", + "DecodedBoxes", + "sigmoid_f32", + "decode_centerhead", + "sort_by_score", + "head_values_at", + "DecodedCells", + "decode_cells", + "DetectedObjects", + "to_detected_objects", +] + +# autoware_perception_msgs/msg/ObjectClassification.msg label values (S:centerpoint:290) +AUTOWARE_LABELS = ("UNKNOWN", "CAR", "TRUCK", "BUS", "TRAILER", "MOTORCYCLE", "BICYCLE", "PEDESTRIAN") +LABEL_IDS = {name: i for i, name in enumerate(AUTOWARE_LABELS)} +# autoware_perception_msgs/msg/DetectedObjectKinematics.msg orientation_availability +UNAVAILABLE, SIGN_UNKNOWN = 0, 1 + +_SEMANTIC = {"CAR": "CAR", "TRUCK": "TRUCK", "BUS": "BUS", "TRAILER": "TRAILER", "BICYCLE": "BICYCLE", + "MOTORBIKE": "MOTORCYCLE", "PEDESTRIAN": "PEDESTRIAN"} + + +def semantic_label(class_name: str) -> int: + """``getSemanticType`` (``ros_utils.cpp:89-108``): network class name -> ObjectClassification label; MOTORBIKE -> + MOTORCYCLE; anything else -> UNKNOWN.""" + return LABEL_IDS[_SEMANTIC.get(class_name, "UNKNOWN")] + + +def is_car_like(label: Any) -> Any: + """``object_recognition_utils::isCarLikeVehicle``: CAR, TRUCK, BUS, TRAILER.""" + lab = np.asarray(label) + return np.isin(lab, [LABEL_IDS["CAR"], LABEL_IDS["TRUCK"], LABEL_IDS["BUS"], LABEL_IDS["TRAILER"]]) + + +def _coerce01(values: Any) -> np.ndarray: + v = np.asarray(values, dtype=np.float32) + return np.where((v >= 0.0) & (v < 1.0), v, np.float32(0.0)).astype(np.float32) + + +@dataclass(frozen=True) +class CenterHeadDecodeConfig: + """Decoder parameters. ``score_thresholds`` is (num_bins, num_classes), the node's ``[bin][class]`` layout + (``node.cpp:129-167``; read as ``thresholds[bin * C + label]``, ``postprocess_kernel.cu:114``).""" + + class_names: tuple + voxel_size_xy: tuple + range_min_xy: tuple + downsample_factor: int + distance_bin_upper_limits: tuple + score_thresholds: np.ndarray + yaw_norm_thresholds: np.ndarray + has_variance: bool = False + has_twist: bool = False + + @classmethod + def create(cls, *, class_names: Sequence[str], voxel_size_xy: Sequence[float], range_min_xy: Sequence[float], + downsample_factor: int, distance_bin_upper_limits: Sequence[float], score_thresholds: Any, + yaw_norm_thresholds: Sequence[float], has_variance: bool = False, + has_twist: bool = False) -> "CenterHeadDecodeConfig": + """Validated config with Autoware's coercions: thresholds outside [0, 1) -> 0, ascending bin limits, one + yaw-norm threshold per class. ``score_thresholds`` may be a scalar, a per-class list, or (bins, classes).""" + names = tuple(class_names) + bins = tuple(float(v) for v in distance_bin_upper_limits) + if list(bins) != sorted(bins): + raise ValueError("distance_bin_upper_limits must be ascending (centerpoint_config.hpp:67-70)") + thr = np.asarray(score_thresholds, dtype=np.float32) + if thr.ndim == 0: + thr = np.full((len(bins), len(names)), thr, np.float32) + elif thr.ndim == 1 and thr.shape[0] == len(names): + thr = np.tile(thr[None, :], (len(bins), 1)) + if thr.shape != (len(bins), len(names)): + raise ValueError(f"score_thresholds must be (bins={len(bins)}, classes={len(names)}), got {thr.shape}") + yaw = np.asarray(yaw_norm_thresholds, dtype=np.float64) + if yaw.shape != (len(names),): + raise ValueError("yaw_norm_thresholds needs one value per class (node.cpp:82-85)") + return cls(names, tuple(float(v) for v in voxel_size_xy), tuple(float(v) for v in range_min_xy), + int(downsample_factor), bins, _coerce01(thr), _coerce01(yaw), bool(has_variance), bool(has_twist)) + + def with_score_threshold(self, value: Optional[float]) -> "CenterHeadDecodeConfig": + """The same config with every class / bin threshold set to ``value`` (coerced like Autoware); None keeps it.""" + if value is None: + return self + return replace(self, score_thresholds=_coerce01(np.full_like(self.score_thresholds, value))) + + @property + def num_classes(self) -> int: + return len(self.class_names) + + +@dataclass +class DecodedBoxes: + """Decoded candidates (struct of arrays, all of length N). ``yaw`` is the network ("mmdet3d") yaw.""" + + cell: np.ndarray # int64 cell index yi * W + xi + label: np.ndarray # int32 class index into class_names + score: np.ndarray # float32 max sigmoid + x: np.ndarray # float32 box centre, model frame + y: np.ndarray + z: np.ndarray + length: np.ndarray # float32 exp(dim[1]) + width: np.ndarray # float32 exp(dim[0]) + height: np.ndarray # float32 exp(dim[2]) + yaw: np.ndarray # float32 atan2(rot0, rot1) + vel_x: np.ndarray + vel_y: np.ndarray + + def __len__(self) -> int: + return int(self.score.shape[0]) + + def take(self, idx: Any) -> "DecodedBoxes": + idx = np.asarray(idx, dtype=np.int64) + return DecodedBoxes(**{f.name: getattr(self, f.name)[idx] for f in fields(self)}) + + def to_dict(self) -> Dict[str, np.ndarray]: + return {f.name: getattr(self, f.name) for f in fields(self)} + + +def sigmoid_f32(x: Any) -> np.ndarray: + """``1.0f / (1.0f + expf(-x))`` in float32 (``postprocess_kernel.cu:46-49``).""" + v = np.asarray(x, dtype=np.float32) + with np.errstate(over="ignore"): + return (np.float32(1.0) / (np.float32(1.0) + np.exp(-v))).astype(np.float32) + + +def decode_centerhead(heads: Mapping[str, Any], cfg: CenterHeadDecodeConfig) -> DecodedBoxes: + """Decode the six head maps ``heatmap (C, H, W)``, ``reg (2, H, W)``, ``height (1, H, W)``, ``dim (3, H, W)``, + ``rot (2, H, W)``, ``vel (2, H, W)`` (a leading batch axis of 1 is accepted) -> candidates in cell order.""" + if cfg.has_variance: + raise NotImplementedError("variance heads (CenterPoint-sigma) are not supported by this decoder") + + def get(name: str, channels: int) -> np.ndarray: + a = np.asarray(heads[name], dtype=np.float32) + if a.ndim == 4: + if a.shape[0] != 1: + raise ValueError(f"{name}: batch {a.shape[0]} != 1") + a = a[0] + if a.ndim != 3 or a.shape[0] < channels: + raise ValueError(f"{name}: expected ({channels}, H, W), got {a.shape}") + return a + + hm = get("heatmap", cfg.num_classes)[:cfg.num_classes] + C, H, W = hm.shape + reg, hei, dim, rot = get("reg", 2), get("height", 1), get("dim", 3), get("rot", 2) + vel = get("vel", 2) if "vel" in heads else np.zeros((2, H, W), np.float32) + for name, a in (("reg", reg), ("height", hei), ("dim", dim), ("rot", rot), ("vel", vel)): + if a.shape[1:] != (H, W): + raise ValueError(f"{name} is {a.shape[1:]}, heatmap is {(H, W)}") + scores = sigmoid_f32(hm) + nan = np.isnan(scores) + if nan.any(): # NaN never wins the strict '>' scan; an all-NaN cell keeps label -1 and score 0 + scores = np.where(nan, np.float32(-np.inf), scores) + label = np.argmax(scores, axis=0).astype(np.int32) # first max == strict '>' from -1 + max_score = np.take_along_axis(scores, label[None].astype(np.int64), 0)[0] + found = np.isfinite(max_score) + max_score = np.where(found, max_score, np.float32(0.0)).astype(np.float32) + yi, xi = np.meshgrid(np.arange(H, dtype=np.float32), np.arange(W, dtype=np.float32), indexing="ij") + f32 = np.float32 + sx = f32(f32(cfg.voxel_size_xy[0]) * f32(cfg.downsample_factor)) + sy = f32(f32(cfg.voxel_size_xy[1]) * f32(cfg.downsample_factor)) + x = (sx * (xi + reg[0]) + f32(cfg.range_min_xy[0])).astype(np.float32) + y = (sy * (yi + reg[1]) + f32(cfg.range_min_xy[1])).astype(np.float32) + radial = np.sqrt(x * x + y * y).astype(np.float32) + bucket = np.full((H, W), -1, dtype=np.int64) + for i in range(len(cfg.distance_bin_upper_limits) - 1, -1, -1): # first upper limit above the distance + bucket = np.where(radial < f32(cfg.distance_bin_upper_limits[i]), i, bucket) + thr = cfg.score_thresholds[np.clip(bucket, 0, None), np.clip(label, 0, None)] + yaw_norm = np.sqrt(rot[0] * rot[0] + rot[1] * rot[1]).astype(np.float32) + yaw_thr = cfg.yaw_norm_thresholds[np.clip(label, 0, None)] + keep = found & (bucket >= 0) & ~(max_score < thr) & (yaw_norm >= yaw_thr) & (max_score > 0.0) + cell = np.nonzero(keep.ravel())[0] + + def at(a: np.ndarray) -> np.ndarray: + return a.reshape(-1)[cell] + + return DecodedBoxes( + cell=cell.astype(np.int64), label=at(label).astype(np.int32), score=at(max_score), x=at(x), y=at(y), + z=at(hei[0]), length=np.exp(at(dim[1])).astype(np.float32), width=np.exp(at(dim[0])).astype(np.float32), + height=np.exp(at(dim[2])).astype(np.float32), yaw=np.arctan2(at(rot[0]), at(rot[1])).astype(np.float32), + vel_x=at(vel[0]), vel_y=at(vel[1])) + + +def sort_by_score(boxes: DecodedBoxes) -> DecodedBoxes: + """Descending score, stable (ties keep cell order): a deterministic ``thrust::sort(..., score_greater())``.""" + return boxes.take(np.argsort(-boxes.score, kind="stable")) + + +# ------------------------------------------------------------------ the decoder at given cells (agreement metrics) + +def head_values_at(heads: Mapping[str, Any], cells: Any) -> Dict[str, np.ndarray]: + """The channels of every head map ``(C, H, W)`` (a leading batch axis of 1 is accepted) at the BEV cells + ``cells`` (``yi * W + xi``, any order) -> ``{name: (C, K) float32}``: an exact, compact form of the maps at the + cells that matter (agreement goldens). :func:`decode_cells` decodes it.""" + idx = np.asarray(cells, dtype=np.int64).reshape(-1) + out: Dict[str, np.ndarray] = {} + for name, a in heads.items(): + a = np.asarray(a, dtype=np.float32) + if a.ndim == 4: + if a.shape[0] != 1: + raise ValueError(f"{name}: batch {a.shape[0]} != 1") + a = a[0] + if a.ndim != 3: + raise ValueError(f"{name}: expected (C, H, W), got {a.shape}") + flat = a.reshape(a.shape[0], -1) + if idx.size and (idx.min() < 0 or idx.max() >= flat.shape[1]): + raise IndexError(f"{name}: cells outside [0, {flat.shape[1]})") + out[name] = np.ascontiguousarray(flat[:, idx]) + return out + + +@dataclass +class DecodedCells: + """:func:`decode_centerhead` evaluated at given cells with **no cell dropped** (struct of arrays, length K). The + box fields are those of :class:`DecodedBoxes`, bit-identical to ``decode_centerhead``'s row for every cell it + keeps (``keep``); the gate inputs and outcomes are kept per cell, so a cell one implementation keeps and another + drops can be explained (score threshold vs yaw-norm gate).""" + + cell: np.ndarray # int64 cell index yi * W + xi + label: np.ndarray # int32 first-max class; -1 when every class score is NaN + score: np.ndarray # float32 max sigmoid (0 when every class score is NaN) + x: np.ndarray + y: np.ndarray + z: np.ndarray + length: np.ndarray + width: np.ndarray + height: np.ndarray + yaw: np.ndarray # float32 network yaw atan2(rot0, rot1) + vel_x: np.ndarray + vel_y: np.ndarray + yaw_norm: np.ndarray # float32 sqrt(rot0^2 + rot1^2): the input of the yaw-norm gate + distance_bin: np.ndarray # int64 first bin with radial distance < its upper limit; -1 beyond the last + score_threshold: np.ndarray # float32 thresholds[bin][label] (bin and label clipped to 0, as the decoder) + yaw_norm_threshold: np.ndarray # float32 yaw-norm threshold of the label + passes_score: np.ndarray # bool: a label, inside a bin, not (score < threshold), score > 0 + passes_yaw_norm: np.ndarray # bool: yaw_norm >= the label's yaw-norm threshold + + @property + def keep(self) -> np.ndarray: + """The cells ``decode_centerhead`` keeps (both gates pass).""" + return self.passes_score & self.passes_yaw_norm + + def __len__(self) -> int: + return int(self.score.shape[0]) + + def take(self, idx: Any) -> "DecodedCells": + idx = np.asarray(idx, dtype=np.int64) + return DecodedCells(**{f.name: getattr(self, f.name)[idx] for f in fields(self)}) + + def to_boxes(self) -> DecodedBoxes: + """The :class:`DecodedBoxes` fields of every cell (gates not applied: ``.take(np.nonzero(keep)[0])`` first + for the decoder's rows).""" + return DecodedBoxes(**{f.name: getattr(self, f.name) for f in fields(DecodedBoxes)}) + + +def decode_cells(values: Mapping[str, Any], cells: Any, grid_w: int, cfg: CenterHeadDecodeConfig) -> DecodedCells: + """The decoder at the cells ``cells`` of a head grid ``grid_w`` cells wide, from their channel values ``values`` + (``{name: (C, K)}``, :func:`head_values_at`; ``vel`` optional), in the float32 arithmetic of + :func:`decode_centerhead` and with the gate outcomes instead of the drop (:class:`DecodedCells`).""" + if cfg.has_variance: + raise NotImplementedError("variance heads (CenterPoint-sigma) are not supported by this decoder") + idx = np.asarray(cells, dtype=np.int64).reshape(-1) + k = idx.size + + def get(name: str, channels: int) -> np.ndarray: + a = np.asarray(values[name], dtype=np.float32) + if a.ndim != 2 or a.shape[0] < channels or a.shape[1] != k: + raise ValueError(f"{name}: expected ({channels}, {k}), got {a.shape}") + return a + + hm = get("heatmap", cfg.num_classes)[:cfg.num_classes] + reg, hei, dim, rot = get("reg", 2), get("height", 1), get("dim", 3), get("rot", 2) + vel = get("vel", 2) if "vel" in values else np.zeros((2, k), np.float32) + scores = sigmoid_f32(hm) + nan = np.isnan(scores) + if nan.any(): + scores = np.where(nan, np.float32(-np.inf), scores) + label = np.argmax(scores, axis=0).astype(np.int32) if k else np.zeros(0, np.int32) + max_score = np.take_along_axis(scores, label[None].astype(np.int64), 0)[0] if k else np.zeros(0, np.float32) + found = np.isfinite(max_score) + max_score = np.where(found, max_score, np.float32(0.0)).astype(np.float32) + f32 = np.float32 + w = int(grid_w) + yi, xi = (idx // w).astype(np.float32), (idx % w).astype(np.float32) + sx = f32(f32(cfg.voxel_size_xy[0]) * f32(cfg.downsample_factor)) + sy = f32(f32(cfg.voxel_size_xy[1]) * f32(cfg.downsample_factor)) + x = (sx * (xi + reg[0]) + f32(cfg.range_min_xy[0])).astype(np.float32) + y = (sy * (yi + reg[1]) + f32(cfg.range_min_xy[1])).astype(np.float32) + radial = np.sqrt(x * x + y * y).astype(np.float32) + bucket = np.full(k, -1, dtype=np.int64) + for i in range(len(cfg.distance_bin_upper_limits) - 1, -1, -1): + bucket = np.where(radial < f32(cfg.distance_bin_upper_limits[i]), i, bucket) + thr = cfg.score_thresholds[np.clip(bucket, 0, None), np.clip(label, 0, None)].astype(np.float32) + yaw_norm = np.sqrt(rot[0] * rot[0] + rot[1] * rot[1]).astype(np.float32) + yaw_thr = cfg.yaw_norm_thresholds[np.clip(label, 0, None)].astype(np.float32) + return DecodedCells( + cell=idx.copy(), label=np.where(found, label, -1).astype(np.int32), score=max_score, x=x, y=y, + z=hei[0].copy(), length=np.exp(dim[1]).astype(np.float32), width=np.exp(dim[0]).astype(np.float32), + height=np.exp(dim[2]).astype(np.float32), yaw=np.arctan2(rot[0], rot[1]).astype(np.float32), + vel_x=vel[0].copy(), vel_y=vel[1].copy(), yaw_norm=yaw_norm, distance_bin=bucket, score_threshold=thr, + yaw_norm_threshold=yaw_thr, passes_score=found & (bucket >= 0) & ~(max_score < thr) & (max_score > 0.0), + passes_yaw_norm=yaw_norm >= yaw_thr) + + +@dataclass +class DetectedObjects: + """Autoware ``DetectedObject`` fields of N objects (struct of arrays). Positions and dimensions hold the float32 + values of the decoder (the message stores them as double); ``yaw`` is the ROS yaw (float32 arithmetic).""" + + label: np.ndarray # uint8 ObjectClassification label + existence_probability: np.ndarray # float32 + x: np.ndarray + y: np.ndarray + z: np.ndarray + yaw: np.ndarray + length: np.ndarray + width: np.ndarray + height: np.ndarray + orientation_availability: np.ndarray # uint8: SIGN_UNKNOWN for car-like labels, else UNAVAILABLE + twist_x: Optional[np.ndarray] = None # object-frame twist (has_twist only) + twist_y: Optional[np.ndarray] = None + source_index: Optional[np.ndarray] = None # row of the DecodedBoxes each object came from + label_before_remap: Optional[np.ndarray] = None + extra: Dict[str, np.ndarray] = field(default_factory=dict) + + def __len__(self) -> int: + return int(self.existence_probability.shape[0]) + + def take(self, idx: Any) -> "DetectedObjects": + idx = np.asarray(idx, dtype=np.int64) + out = {} + for f in fields(self): + v = getattr(self, f.name) + if f.name == "extra": + out[f.name] = {k: a[idx] for k, a in v.items()} + else: + out[f.name] = None if v is None else v[idx] + return DetectedObjects(**out) + + def boxes_xyzlwh_yaw(self) -> np.ndarray: + """(N, 7) float32 x, y, z, length, width, height, yaw (the ``ttaw.outputs.Detections3D`` layout).""" + return np.stack([self.x, self.y, self.z, self.length, self.width, self.height, self.yaw], + axis=1).astype(np.float32) + + def to_records(self, labels: Sequence[str] = AUTOWARE_LABELS) -> List[Dict[str, Any]]: + out = [] + for i in range(len(self)): + d = {"label": labels[int(self.label[i])], "existence_probability": float(self.existence_probability[i]), + "x": float(self.x[i]), "y": float(self.y[i]), "z": float(self.z[i]), "yaw": float(self.yaw[i]), + "length": float(self.length[i]), "width": float(self.width[i]), "height": float(self.height[i]), + "orientation_availability": int(self.orientation_availability[i])} + if self.label_before_remap is not None and self.label_before_remap[i] != self.label[i]: + d["label_before_remap"] = labels[int(self.label_before_remap[i])] + out.append(d) + return out + + +def to_detected_objects(boxes: DecodedBoxes, class_names: Sequence[str], *, has_twist: bool = False) -> DetectedObjects: + """``box3DToDetectedObject`` for every row (module docstring).""" + table = np.array([semantic_label(n) for n in class_names], dtype=np.uint8) + lab = np.asarray(boxes.label, dtype=np.int64) + valid = (lab >= 0) & (lab < len(class_names)) + label = np.where(valid, table[np.clip(lab, 0, len(class_names) - 1)], LABEL_IDS["UNKNOWN"]).astype(np.uint8) + yaw = (-np.asarray(boxes.yaw, dtype=np.float32).astype(np.float64) - math.pi / 2.0).astype(np.float32) + orient = np.where(is_car_like(label), SIGN_UNKNOWN, UNAVAILABLE).astype(np.uint8) + tx = ty = None + if has_twist: + c, s = np.cos(yaw), np.sin(yaw) # float32, as std::cos(float) + tx = (c * boxes.vel_x + s * boxes.vel_y).astype(np.float32) + ty = (-s * boxes.vel_x + c * boxes.vel_y).astype(np.float32) + return DetectedObjects(label=label, existence_probability=np.asarray(boxes.score, np.float32), + x=boxes.x.copy(), y=boxes.y.copy(), z=boxes.z.copy(), yaw=yaw, + length=boxes.length.copy(), width=boxes.width.copy(), height=boxes.height.copy(), + orientation_availability=orient, twist_x=tx, twist_y=ty, + source_index=np.arange(len(boxes), dtype=np.int64), label_before_remap=label.copy(), + extra={"cell": np.asarray(boxes.cell, dtype=np.int64)}) diff --git a/code/tt_diffusion_planner/ttaw/device.py b/code/tt_diffusion_planner/ttaw/device.py new file mode 100644 index 0000000000000000000000000000000000000000..426cecd8a53a00d408c421c5a6ab0568c53f4fc5 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/device.py @@ -0,0 +1,309 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C01: open the Blackhole p150 the way every published number is measured. + +Default: dispatch on the idle Ethernet cores (``ttnn.DispatchCoreConfig(ttnn.DispatchCoreType.ETH)``), which frees the +Tensix dispatch column and gives a 12x10 = 120-core compute grid on a p150b (11x10 with stock WORKER dispatch). ETH +dispatch needs ``patches/tt-metal-eth-dispatch.patch`` on tt-metal 44d6650; the patch adds the marker +``single_chip_arch_1cq_no_dispatch_s`` to ``tt_metal/impl/dispatch/topology.cpp``, which ``dispatch="auto"`` looks +for. If the ETH open fails, a ``RuntimeWarning`` is issued and WORKER dispatch is used (``allow_fallback=False`` +raises instead, for the container smoke). Never hard-code the grid: use :func:`compute_grid` / +``device.compute_with_storage_grid_size()``. + +``ttnn`` is imported inside the functions, so importing this module has no side effects. + +Example:: + + from .ttaw.device import device_session, describe_device + + with device_session(dispatch="eth", num_command_queues=2, trace_region_size=64 << 20) as dev: + print(describe_device(dev)) # {'dispatch': 'eth', 'grid': '12x10', 'cores': 120, 'num_command_queues': 2, ...} +""" +from __future__ import annotations + +import contextlib +import importlib.util +import os +import warnings +from dataclasses import asdict, dataclass +from pathlib import Path +from typing import Any, Dict, Iterator, Mapping, Optional, Tuple + +__all__ = [ + "DISPATCH_MODES", + "PATCH_MARKER", + "DEFAULT_L1_SMALL_SIZE", + "DEFAULT_TRACE_REGION_SIZE", + "P150_ETH_GRID", + "DeviceConfig", + "tt_metal_home", + "eth_dispatch_patch_present", + "resolve_dispatch", + "open_device", + "describe_device", + "open_info", + "close_device", + "device_session", + "compute_grid", + "core_grid", + "full_core_range_set", +] + +DISPATCH_MODES = ("eth", "worker", "auto") +PATCH_MARKER = "single_chip_arch_1cq_no_dispatch_s" +PATCH_RELPATH = Path("tt_metal", "impl", "dispatch", "topology.cpp") +DEFAULT_L1_SMALL_SIZE = 32768 # conv / pool config tensors (CNN demos use 24576-32768) +DEFAULT_TRACE_REGION_SIZE = 64 << 20 # DRAM bytes for the trace command buffers of all captured traces +P150_ETH_GRID = (12, 10) # p150b compute grid with ETH dispatch (research/DISPATCH.md) + +# id(device) -> (device, what open_device actually did: dispatch used, CQs, sizes); read by describe_device. The +# device itself is kept so that its id cannot be reused by another object while the entry exists (a device opened +# outside ttaw must never inherit a dead device's record); close_device removes the entry. +_OPEN_INFO: Dict[int, Tuple[Any, Dict[str, Any]]] = {} +# id(device) -> device for every device closed by close_device: a second close is a no-op. The reference keeps the +# closed object alive, so its id cannot be reused by a later device (ttnn device objects take no attributes). +_CLOSED: Dict[int, Any] = {} + + +def _normalize_dispatch(dispatch: Optional[str]) -> str: + mode = (dispatch or "eth").strip().lower() + if mode not in DISPATCH_MODES: + raise ValueError(f"dispatch={dispatch!r}: expected one of {DISPATCH_MODES}") + return mode + + +@dataclass(frozen=True) +class DeviceConfig: + """Everything :func:`open_device` needs, so a model can pin it, log it and A/B it from the environment.""" + + device_id: int = 0 + dispatch: str = "eth" + num_command_queues: int = 1 + l1_small_size: int = DEFAULT_L1_SMALL_SIZE + trace_region_size: int = DEFAULT_TRACE_REGION_SIZE + worker_l1_size: Optional[int] = None + allow_fallback: bool = True + + def __post_init__(self) -> None: + object.__setattr__(self, "dispatch", _normalize_dispatch(self.dispatch)) + if self.num_command_queues not in (1, 2): + raise ValueError(f"num_command_queues={self.num_command_queues}: expected 1 or 2") + if self.l1_small_size < 0 or self.trace_region_size < 0: + raise ValueError("l1_small_size and trace_region_size must be >= 0") + if self.worker_l1_size is not None and self.worker_l1_size <= 0: + raise ValueError("worker_l1_size must be a positive byte count (or None for tt-metal's default)") + + @classmethod + def from_env(cls, prefix: str, env: Optional[Mapping[str, str]] = None, **defaults: Any) -> "DeviceConfig": + """Model defaults overridden by ``_DISPATCH`` (eth|worker|auto), ``_NUM_CQS``, + ``_L1_SMALL``, ``_TRACE_REGION``, ``_WORKER_L1_SIZE`` and ``TT_DEVICE_ID`` + (BUNDLE_CONVENTIONS.md section 7.6). Empty or ``0`` values keep the default. Read once, at build.""" + env = os.environ if env is None else env + values = dict(defaults) + + def _int(name: str) -> Optional[int]: + raw = (env.get(name) or "").strip() + if not raw: + return None + try: + value = int(float(raw)) + except ValueError: + raise ValueError(f"{name}={raw!r} is not an integer") from None + return value or None + + dispatch = (env.get(f"{prefix}_DISPATCH") or "").strip() + if dispatch: + values["dispatch"] = dispatch + for key, name in (("num_command_queues", f"{prefix}_NUM_CQS"), ("l1_small_size", f"{prefix}_L1_SMALL"), + ("trace_region_size", f"{prefix}_TRACE_REGION"), + ("worker_l1_size", f"{prefix}_WORKER_L1_SIZE")): + value = _int(name) + if value is not None: + values[key] = value + device_id = (env.get("TT_DEVICE_ID") or "").strip() + if device_id: + values["device_id"] = int(device_id) + return cls(**values) + + def open(self): + """``open_device(**self)``.""" + return open_device(**asdict(self)) + + +def tt_metal_home() -> Optional[Path]: + """The tt-metal tree: ``$TT_METAL_HOME``, else ``$TT_METAL_RUNTIME_ROOT``, else the tree holding the ``ttnn`` + package (located with ``importlib.util.find_spec``, so ttnn is not imported).""" + for var in ("TT_METAL_HOME", "TT_METAL_RUNTIME_ROOT"): + value = os.environ.get(var) + if value and Path(value).is_dir(): + return Path(value) + try: + spec = importlib.util.find_spec("ttnn") + except (ImportError, ValueError): + spec = None + if spec is None or not spec.origin: + return None + for parent in Path(spec.origin).resolve().parents: + if (parent / "tt_metal").is_dir(): + return parent + return None + + +def eth_dispatch_patch_present(home: Optional[os.PathLike] = None) -> bool: + """True when the tt-metal tree carries the ETH-dispatch patch (its marker string is in ``topology.cpp``).""" + root = tt_metal_home() if home is None else Path(home) + if root is None: + return False + try: + return PATCH_MARKER in (Path(root) / PATCH_RELPATH).read_text(errors="ignore") + except OSError: + return False + + +def resolve_dispatch(dispatch: str = "auto", *, home: Optional[os.PathLike] = None) -> str: + """``"eth"`` / ``"worker"`` pass through; ``"auto"`` becomes ``"eth"`` when the patch marker is found, else + ``"worker"`` with a ``RuntimeWarning`` (the grid is then 11x10 and the published numbers do not apply).""" + mode = _normalize_dispatch(dispatch) + if mode != "auto": + return mode + if eth_dispatch_patch_present(home): + return "eth" + warnings.warn( + "the tt-metal ETH-dispatch patch was not found (marker " + f"{PATCH_MARKER!r} missing from {PATCH_RELPATH} under {tt_metal_home()}); using WORKER dispatch: the compute " + "grid is 11x10 on a p150 and the published numbers do not apply (apply patches/tt-metal-eth-dispatch.patch)", + RuntimeWarning, stacklevel=3) + return "worker" + + +def open_device(device_id: int = 0, *, dispatch: str = "eth", num_command_queues: int = 1, + l1_small_size: int = DEFAULT_L1_SMALL_SIZE, trace_region_size: int = DEFAULT_TRACE_REGION_SIZE, + worker_l1_size: Optional[int] = None, allow_fallback: bool = True): + """Open one chip and return the ttnn (1x1 mesh) device. + + Args: + device_id: chip id as ttnn sees it. + dispatch: ``"eth"`` (default; 12x10 on a p150b), ``"worker"`` (A/B only; 11x10) or ``"auto"`` (ETH when the + patch marker is present, else WORKER with a ``RuntimeWarning``). + num_command_queues: 1, or 2 for trace + 2CQ (input upload on CQ1). + l1_small_size: L1_SMALL bytes per core (conv / pool config tensors). + trace_region_size: DRAM bytes reserved for trace buffers (0 = dynamic top-down buffers). + worker_l1_size: allocatable L1 per core in bytes; smaller values grow the kernel-config ring buffer + (TT_PLATFORM.md section 1). ``None`` keeps tt-metal's default. + allow_fallback: when the ETH open fails, warn and open with WORKER dispatch (``True``) or re-raise. + """ + import ttnn + + cfg = DeviceConfig(device_id=device_id, dispatch=dispatch, num_command_queues=num_command_queues, + l1_small_size=l1_small_size, trace_region_size=trace_region_size, + worker_l1_size=worker_l1_size, allow_fallback=allow_fallback) + mode = resolve_dispatch(cfg.dispatch) + params: Dict[str, Any] = dict(l1_small_size=cfg.l1_small_size, trace_region_size=cfg.trace_region_size, + num_command_queues=cfg.num_command_queues) + if cfg.worker_l1_size is not None: + params["worker_l1_size"] = cfg.worker_l1_size + device = None + fallback: Optional[str] = None + if mode == "eth": + try: + device = ttnn.open_device(device_id=cfg.device_id, + dispatch_core_config=ttnn.DispatchCoreConfig(ttnn.DispatchCoreType.ETH), **params) + except Exception as exc: # noqa: BLE001 -- tt-metal without the ETH-dispatch patch, or no idle ETH cores + if not cfg.allow_fallback: + raise + fallback = f"{type(exc).__name__}: {exc}" + warnings.warn(f"ETH dispatch is not available ({fallback}); falling back to WORKER dispatch: the compute " + "grid is 11x10 on a p150 and the published numbers do not apply " + "(apply patches/tt-metal-eth-dispatch.patch)", RuntimeWarning, stacklevel=2) + mode = "worker" + if device is None: + device = ttnn.open_device(device_id=cfg.device_id, + dispatch_core_config=ttnn.DispatchCoreConfig(ttnn.DispatchCoreType.WORKER), **params) + _CLOSED.pop(id(device), None) + _OPEN_INFO[id(device)] = (device, { + "dispatch": mode, "dispatch_requested": cfg.dispatch, "fallback": fallback, "device_id": cfg.device_id, + "num_command_queues": cfg.num_command_queues, "l1_small_size": cfg.l1_small_size, + "trace_region_size": cfg.trace_region_size, "worker_l1_size": cfg.worker_l1_size, + }) + gx, gy = compute_grid(device) + if mode == "eth" and (gx, gy) != P150_ETH_GRID: + warnings.warn(f"compute grid is {gx}x{gy}, expected {P150_ETH_GRID[0]}x{P150_ETH_GRID[1]} with ETH dispatch " + "on a p150b", RuntimeWarning, stacklevel=2) + return device + + +def compute_grid(device) -> Tuple[int, int]: + """``(x, y)`` of ``device.compute_with_storage_grid_size()``: (12, 10) with ETH dispatch on a p150b.""" + g = device.compute_with_storage_grid_size() + return int(g.x), int(g.y) + + +def core_grid(device): + """``ttnn.CoreGrid`` covering the whole compute grid (for ``core_grid=`` / program configs).""" + import ttnn + + x, y = compute_grid(device) + return ttnn.CoreGrid(y=y, x=x) + + +def full_core_range_set(device): + """``ttnn.CoreRangeSet`` of one rectangle covering the whole compute grid (for ``generic_op`` descriptors).""" + import ttnn + + x, y = compute_grid(device) + return ttnn.CoreRangeSet([ttnn.CoreRange(ttnn.CoreCoord(0, 0), ttnn.CoreCoord(x - 1, y - 1))]) + + +def open_info(device) -> Dict[str, Any]: + """What :func:`open_device` recorded for ``device`` (dispatch used, CQs, sizes), ``{}`` for a device opened + elsewhere.""" + entry = _OPEN_INFO.get(id(device)) + return dict(entry[1]) if entry is not None and entry[0] is device else {} + + +def _arch_name(device) -> str: + try: + arch = device.arch() + except Exception: # noqa: BLE001 -- a fake or partially closed device + return "unknown" + name = getattr(arch, "name", None) or str(arch).rsplit(".", 1)[-1] + return name.lower() + + +def describe_device(device, device_id: Optional[int] = None) -> Dict[str, Any]: + """What ``/info`` and ``model.info`` report: dispatch actually used, CQs, grid, cores, arch and open parameters. + + A device opened outside :func:`open_device` reports ``dispatch="unknown"``.""" + gx, gy = compute_grid(device) + info: Dict[str, Any] = {"dispatch": "unknown", "num_command_queues": None} + info.update(open_info(device)) + info.update({"grid": f"{gx}x{gy}", "grid_x": gx, "grid_y": gy, "cores": gx * gy, "arch": _arch_name(device), + "eth_patch": eth_dispatch_patch_present()}) + if device_id is not None: + info["device_id"] = device_id + return info + + +def close_device(device) -> None: + """Close a device (synchronizes first, as ``ttnn.close_device`` does). Idempotent: closing a device this + function already closed does nothing.""" + import ttnn + + if _CLOSED.get(id(device)) is device: + return + if open_info(device): + del _OPEN_INFO[id(device)] + try: + ttnn.close_device(device) + finally: + _CLOSED[id(device)] = device + + +@contextlib.contextmanager +def device_session(device_id: int = 0, **kwargs: Any) -> Iterator[Any]: + """``with device_session(dispatch="eth", num_command_queues=2) as dev:`` -- opens with :func:`open_device` + (same keyword arguments) and always closes, also when the body raises.""" + device = open_device(device_id, **kwargs) + try: + yield device + finally: + close_device(device) diff --git a/code/tt_diffusion_planner/ttaw/geometry.py b/code/tt_diffusion_planner/ttaw/geometry.py new file mode 100644 index 0000000000000000000000000000000000000000..bee2099f90087e7c1d6f04da5b9e41a3ca7e6ae7 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/geometry.py @@ -0,0 +1,194 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C10 geometry (sweep and ego-motion parts): rigid transforms in float64, Autoware's float32 sweep transform, and +the yaw / quaternion conventions of the LiDAR detectors. + +Naming: ``T_a_from_b`` maps points of frame ``b`` into frame ``a`` (4x4, metres, row-major, translation in the last +column), so ``p_a = T_a_from_b @ p_b`` and ``T_a_from_c = T_a_from_b @ T_b_from_c``. + +Two precisions on purpose: + +- **float64** for composing poses (``compose``, ``invert_rigid``): what tf2 does, and what every port uses unless its + Autoware node does otherwise (S:streampetr:386; S:pointpainting:212-215; S:bevformer:437). +- **float32 as Autoware's lidar_centerpoint does it**: the node casts each TF to ``Eigen::Affine3f``, inverts and + multiplies in float32 (``pointcloud_densification.cpp:47-51,82-86``; ``voxel_generator.cpp:63``) and the CUDA + kernel ``generateSweepPoints_kernel`` applies the float32 affine as ``m0*x + m4*y + m8*z + m12`` (column-major, + ``preprocess_kernel.cu:58-90``). :func:`invert_affine_f32` and :func:`transform_points_f32` reproduce that + arithmetic in numpy float32 without fused multiply-adds, so they agree with the GPU to the last bit except where + nvcc contracts a multiply-add (an FMA rounds once): a sub-micrometre difference, irrelevant to 0.32 m pillars. + +Yaw conventions (S:centerpoint:305-307): the deployed CenterPoint ONNX regresses the "mmdet3d" yaw; Autoware +publishes ``yaw_ros = -yaw_net - pi/2`` counter-clockwise from +x of ``base_link`` (``ros_utils.cpp:55-60``). + +numpy only; no side effects on import. +""" +from __future__ import annotations + +import math +from typing import Any, Sequence + +import numpy as np + +from .io import parse_transform + +__all__ = [ + "as_transform", + "compose", + "invert_rigid", + "invert_affine_f32", + "compose_f32", + "transform_points", + "transform_points_f32", + "rotation_from_rpy", + "transform_from_xyz_rpy", + "quaternion_wxyz_from_yaw", + "yaw_from_quaternion_wxyz", + "yaw_from_rotation", + "wrap_angle", + "mmdet_yaw_to_ros", + "ros_yaw_to_mmdet", +] + + +# ------------------------------------------------------------------------------------------------- transforms + +def as_transform(obj: Any, *, field: str = "transform") -> np.ndarray: + """Any spelling :func:`ttaw.io.parse_transform` accepts (4x4 / 3x4 array, ``{matrix}``, + ``{translation, rotation_wxyz | rotation_xyzw | rotation}``, ``{x, y, z, roll, pitch, yaw}``) -> a validated + (4, 4) float64 rigid transform. A malformed value raises :class:`ttaw.io.InputError`.""" + return parse_transform(obj, field=field) + + +def compose(*transforms: np.ndarray) -> np.ndarray: + """``T_0 @ T_1 @ ... @ T_n`` in float64 (``compose(T_a_from_b, T_b_from_c) == T_a_from_c``).""" + if not transforms: + return np.eye(4) + out = np.asarray(transforms[0], dtype=np.float64) + for t in transforms[1:]: + out = out @ np.asarray(t, dtype=np.float64) + return out + + +def invert_rigid(T: np.ndarray) -> np.ndarray: + """Exact inverse of a rigid transform in float64: ``[R^T, -R^T t]``.""" + T = np.asarray(T, dtype=np.float64) + out = np.eye(4) + r = T[:3, :3] + out[:3, :3] = r.T + out[:3, 3] = -(r.T @ T[:3, 3]) + return out + + +def invert_affine_f32(T: np.ndarray) -> np.ndarray: + """``Eigen::Affine3f::inverse()`` (TransformTraits ``Affine``) in float32: the 3x3 linear part inverted through + its cofactors and the column-0 determinant (Eigen's ``compute_inverse<..., 3>``), translation ``-L^-1 t``, + every operation rounded to float32. Used by the Autoware densification emulation + (``pointcloud_densification.cpp:84``).""" + m = np.asarray(T, dtype=np.float32) + a = m[:3, :3] + f = np.float32 + + def cofactor(i: int, j: int) -> np.float32: # Eigen cofactor_3x3 + i1, i2, j1, j2 = (i + 1) % 3, (i + 2) % 3, (j + 1) % 3, (j + 2) % 3 + return f(f(a[i1, j1] * a[i2, j2]) - f(a[i1, j2] * a[i2, j1])) + + det = f(f(cofactor(0, 0) * a[0, 0]) + f(cofactor(1, 0) * a[1, 0])) + f(cofactor(2, 0) * a[2, 0]) + if det == 0 or not np.isfinite(det): + raise ValueError("singular affine transform") + inv_det = f(f(1.0) / det) + inv = np.empty((3, 3), np.float32) + for i in range(3): + for j in range(3): + inv[i, j] = f(cofactor(j, i) * inv_det) # adjugate / det + out = np.eye(4, dtype=np.float32) + out[:3, :3] = inv + t = m[:3, 3] + for i in range(3): + out[i, 3] = -(f(f(inv[i, 0] * t[0]) + f(inv[i, 1] * t[1])) + f(inv[i, 2] * t[2])) + return out + + +def compose_f32(A: np.ndarray, B: np.ndarray) -> np.ndarray: + """``A @ B`` of two float32 affine matrices with float32 rounding after every product and sum, summed in index + order (Eigen's lazy product of two ``Affine3f``).""" + a = np.asarray(A, dtype=np.float32) + b = np.asarray(B, dtype=np.float32) + out = np.zeros((4, 4), np.float32) + for k in range(4): + out = (out + a[:, k:k + 1] * b[k:k + 1, :]).astype(np.float32) + return out + + +def transform_points(points: np.ndarray, T: np.ndarray) -> np.ndarray: + """``T @ [x, y, z, 1]`` for (N, >=3) points in float64 (extra columns are passed through unchanged).""" + p = np.asarray(points, dtype=np.float64) + T = np.asarray(T, dtype=np.float64) + out = p.copy() + out[:, :3] = p[:, :3] @ T[:3, :3].T + T[:3, 3] + return out + + +def transform_points_f32(xyz: np.ndarray, T: np.ndarray) -> np.ndarray: + """Autoware ``generateSweepPoints_kernel`` (``preprocess_kernel.cu:58-90``): the affine is cast to float32 and + each output coordinate is ``((m[r,0]*x + m[r,1]*y) + m[r,2]*z) + m[r,3]`` in float32, left to right (no FMA). + ``xyz``: (N, >=3); returns (N, 3) float32.""" + p = np.asarray(xyz, dtype=np.float32) + m = np.asarray(T, dtype=np.float32) + x, y, z = p[:, 0], p[:, 1], p[:, 2] + out = np.empty((len(p), 3), np.float32) + for r in range(3): + out[:, r] = ((m[r, 0] * x + m[r, 1] * y) + m[r, 2] * z) + m[r, 3] + return out + + +# ----------------------------------------------------------------------------------------------- rotations + +def rotation_from_rpy(roll: float, pitch: float, yaw: float) -> np.ndarray: + """tf2 ``setRPY``: ``R = Rz(yaw) @ Ry(pitch) @ Rx(roll)`` (float64).""" + cr, sr, cp, sp, cy, sy = (math.cos(roll), math.sin(roll), math.cos(pitch), math.sin(pitch), math.cos(yaw), + math.sin(yaw)) + return np.array([[cy * cp, cy * sp * sr - sy * cr, cy * sp * cr + sy * sr], + [sy * cp, sy * sp * sr + cy * cr, sy * sp * cr - cy * sr], + [-sp, cp * sr, cp * cr]], dtype=np.float64) + + +def transform_from_xyz_rpy(x: float, y: float, z: float, roll: float = 0.0, pitch: float = 0.0, + yaw: float = 0.0) -> np.ndarray: + """A (4, 4) float64 transform from a translation and tf2 RPY angles (Autoware's calibration yaml spelling).""" + out = np.eye(4) + out[:3, :3] = rotation_from_rpy(roll, pitch, yaw) + out[:3, 3] = (x, y, z) + return out + + +def quaternion_wxyz_from_yaw(yaw: float) -> np.ndarray: + """``autoware_utils::create_quaternion_from_yaw``: rotation about +z, as (w, x, y, z) float64.""" + return np.array([math.cos(yaw / 2.0), 0.0, 0.0, math.sin(yaw / 2.0)], dtype=np.float64) + + +def yaw_from_quaternion_wxyz(q: Sequence[float]) -> float: + """``tf2::getYaw`` of a (w, x, y, z) quaternion.""" + w, x, y, z = (float(v) for v in q) + return math.atan2(2.0 * (w * z + x * y), 1.0 - 2.0 * (y * y + z * z)) + + +def yaw_from_rotation(r: np.ndarray) -> float: + """Heading of a rotation matrix (rotation of +x about +z, ``atan2(R[1,0], R[0,0])``).""" + r = np.asarray(r, dtype=np.float64) + return math.atan2(r[1, 0], r[0, 0]) + + +def wrap_angle(a: Any) -> Any: + """Wrap angles to [-pi, pi).""" + return (np.asarray(a, dtype=np.float64) + math.pi) % (2.0 * math.pi) - math.pi + + +def mmdet_yaw_to_ros(yaw_net: Any) -> np.ndarray: + """``ros_utils.cpp:55``: ``const float yaw = -box3d.yaw - pi / 2`` (float input, double arithmetic, float + result). Returns float32 radians, not wrapped (as Autoware).""" + y = np.asarray(yaw_net, dtype=np.float32).astype(np.float64) + return (-y - math.pi / 2.0).astype(np.float32) + + +def ros_yaw_to_mmdet(yaw_ros: Any) -> np.ndarray: + """Inverse of :func:`mmdet_yaw_to_ros` (float64): ``yaw_net = -yaw_ros - pi/2``.""" + return -np.asarray(yaw_ros, dtype=np.float64) - math.pi / 2.0 diff --git a/code/tt_diffusion_planner/ttaw/golden.py b/code/tt_diffusion_planner/ttaw/golden.py new file mode 100644 index 0000000000000000000000000000000000000000..e00fb389b91d619962eb1fae59d72751ed7aec6a --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/golden.py @@ -0,0 +1,488 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C06 goldens: golden tensor files, a tap registry, gates that are never loosened, JSON compare reports. + +- **Golden files** are ``.npz`` (loaded with ``allow_pickle=False``) holding named tensors ("taps") and an optional + JSON metadata record under the key ``__meta__`` (source file hashes, reference script, sample, date). +- **Taps** are named intermediates recorded by :class:`TapRegistry` from the CPU reference or from an *eager* + device run (never inside a trace capture: a traced intermediate is freed at the end of the capture). +- **Gates** (PLAN.md section 0.3): thresholds are declared in the test (from the spec targets) and frozen in a JSON + file next to the test at the first green run (``test_pcc_device.py`` -> ``test_pcc_device.gates.json``). From then + on a declaration looser than the frozen threshold raises :class:`GateLoosenedError`; tightening is allowed and + re-frozen. Changing a frozen gate means editing the JSON by hand and disclosing it (VERIFICATION_.md). +- **Reports**: :class:`CompareReport` collects per-tap statistics and gate verdicts and writes JSON. + +numpy only (device tensors are converted through :func:`.tensors.to_numpy`, which imports ttnn lazily). +""" +from __future__ import annotations + +import contextlib +import datetime as _dt +import fnmatch +import json +import math +import os +import sys +from dataclasses import asdict, dataclass, field +from pathlib import Path +from typing import Any, Dict, Iterator, List, Mapping, Optional, Sequence, Union + +import numpy as np + +from . import metrics as _m +from .weights import file_sha256 + +__all__ = [ + "save_goldens", + "load_goldens", + "Goldens", + "TapRegistry", + "Gate", + "GateResult", + "GateLoosenedError", + "GateRegistry", + "METRICS", + "DIRECTIONS", + "TapComparison", + "CompareReport", + "compare_taps", +] + +PathLike = Union[str, os.PathLike] +META_KEY = "__meta__" + + +def _now() -> str: + return _dt.datetime.now(_dt.timezone.utc).isoformat(timespec="seconds") + + +def _host(value: Any) -> np.ndarray: + from .tensors import to_numpy + + return np.asarray(to_numpy(value)) + + +def _atomic_write_text(path: Path, text: str) -> None: + path.parent.mkdir(parents=True, exist_ok=True) + tmp = path.with_name(f".{path.name}.{os.getpid()}.tmp") + tmp.write_text(text) + os.replace(tmp, path) + + +# --------------------------------------------------------------------------------------------- golden files + +def save_goldens(path: PathLike, tensors: Mapping[str, Any], meta: Optional[Mapping[str, Any]] = None, *, + compress: bool = False) -> Path: + """Write ``tensors`` (numpy / torch / ttnn) and ``meta`` (JSON-able) to an ``.npz`` file.""" + path = Path(path) + arrays = {str(k): _host(v) for k, v in tensors.items()} + if META_KEY in arrays: + raise ValueError(f"{META_KEY!r} is reserved for metadata") + arrays[META_KEY] = np.array(json.dumps(dict(meta or {}), sort_keys=True, default=str)) + path.parent.mkdir(parents=True, exist_ok=True) + tmp = path.with_name(f".{path.stem}.{os.getpid()}.tmp.npz") + (np.savez_compressed if compress else np.savez)(tmp, **arrays) + os.replace(tmp, path) + return path + + +class Goldens(Mapping): + """Read-only, lazily loaded golden tensors of one ``.npz`` file (``with load_goldens(p) as g: g["neck"]``).""" + + def __init__(self, path: PathLike): + self.path = Path(path) + self._npz = np.load(self.path, allow_pickle=False) + self._cache: Dict[str, np.ndarray] = {} + self.meta: Dict[str, Any] = (json.loads(str(self._npz[META_KEY])) if META_KEY in self._npz.files else {}) + self._sha256: Optional[str] = None + + @property + def sha256(self) -> str: + if self._sha256 is None: + self._sha256 = file_sha256(self.path) + return self._sha256 + + def __getitem__(self, name: str) -> np.ndarray: + if name == META_KEY or name not in self._npz.files: + raise KeyError(f"{name!r} not in {self.path.name}; have {sorted(self)[:20]}") + if name not in self._cache: + self._cache[name] = self._npz[name] + return self._cache[name] + + def __iter__(self) -> Iterator[str]: + return (k for k in self._npz.files if k != META_KEY) + + def __len__(self) -> int: + return sum(1 for _ in self) + + def close(self) -> None: + self._npz.close() + + def __enter__(self) -> "Goldens": + return self + + def __exit__(self, *exc) -> None: + self.close() + + +def load_goldens(path: PathLike) -> Goldens: + return Goldens(path) + + +# ------------------------------------------------------------------------------------------------- taps + +class TapRegistry: + """Records named intermediates. ``tap(name, value)`` returns ``value`` unchanged, so it can be inlined:: + + taps = TapRegistry(include=["backbone.*", "neck"]) + x = taps.tap("backbone.layer1", layer1(x)) # stored as numpy (host read for device tensors) + with taps.scope("head"): + y = taps.tap("cls", cls_head(x)) # recorded as "head.cls" + taps.save("golden/sample0_taps.npz", meta={...}) + + A disabled registry (``TapRegistry(enabled=False)``, or :data:`NULL_TAPS`) does nothing, so production code can + keep its ``tap`` calls. Device tensors are read immediately (they may be deallocated later by the model); + calling ``tap`` on a device tensor while a trace capture is active raises.""" + + def __init__(self, enabled: bool = True, *, include: Sequence[str] = ("*",), exclude: Sequence[str] = ()): + self.enabled = enabled + self.include = tuple(include) + self.exclude = tuple(exclude) + self._prefix: List[str] = [] + self._values: Dict[str, np.ndarray] = {} + + def _full(self, name: str) -> str: + return ".".join(self._prefix + [name]) + + def wants(self, name: str) -> bool: + """Whether ``name`` (full dotted name) would be recorded.""" + return (self.enabled and any(fnmatch.fnmatchcase(name, p) for p in self.include) + and not any(fnmatch.fnmatchcase(name, p) for p in self.exclude)) + + def tap(self, name: str, value: Any) -> Any: + if not self.enabled: + return value + full = self._full(name) + if not self.wants(full): + return value + ttnn = sys.modules.get("ttnn") + if ttnn is not None and isinstance(value, getattr(ttnn, "Tensor", ())): + if value.storage_type() == ttnn.StorageType.DEVICE and ttnn.is_trace_capture_active(value.device()): + raise RuntimeError(f"tap {full!r} inside a trace capture: return it as a trace output instead") + self._values[full] = _host(value) + return value + + @contextlib.contextmanager + def scope(self, name: str) -> Iterator["TapRegistry"]: + self._prefix.append(name) + try: + yield self + finally: + self._prefix.pop() + + def names(self) -> List[str]: + return list(self._values) + + def __contains__(self, name: str) -> bool: + return name in self._values + + def __getitem__(self, name: str) -> np.ndarray: + return self._values[name] + + def items(self): + return self._values.items() + + def to_dict(self) -> Dict[str, np.ndarray]: + return dict(self._values) + + def clear(self) -> None: + self._values.clear() + + def save(self, path: PathLike, meta: Optional[Mapping[str, Any]] = None, *, compress: bool = False) -> Path: + return save_goldens(path, self._values, meta, compress=compress) + + +NULL_TAPS = TapRegistry(enabled=False) + +# ------------------------------------------------------------------------------------------------ gates + +# metric name -> (function(test, ref, **kw) -> float, direction): "min" = higher is better (value >= threshold) +METRICS: Dict[str, Any] = { + "pcc": (_m.pcc, "min"), + "masked_pcc": (_m.masked_pcc, "min"), + "argmax_agreement": (_m.argmax_agreement, "min"), + "label_agreement": (_m.label_agreement, "min"), + "mask_iou": (_m.mask_iou, "min"), + "max_abs": (lambda t, r: _m.error_stats(t, r)["max_abs"], "max"), + "mean_abs": (lambda t, r: _m.error_stats(t, r)["mean_abs"], "max"), + "rel_l2": (lambda t, r: _m.error_stats(t, r)["rel_l2"], "max"), +} + +# The direction of every metric name a gate may use without spelling it out (the METRICS above plus the scalars +# recorded with CompareReport.add_value). A bare-number gate on any other metric is refused rather than guessed: +# a guessed "min" on an error metric (ADE, max |error|) would pass exactly the results it is meant to fail. +DIRECTIONS: Dict[str, str] = { + **{name: direction for name, (_, direction) in METRICS.items()}, + **dict.fromkeys(("valid_row_pcc", "agreement", "iou", "miou", "mean_iou", "recall", "precision", "topk_overlap", + "topk_set_overlap"), "min"), + **dict.fromkeys(("max_abs_err", "ade", "fde", "center_err", "center_err_max", "center_err_mean", "score_err", + "score_err_max", "score_err_mean"), "max"), +} + + +@dataclass(frozen=True) +class Gate: + """``direction="min"``: pass when ``value >= threshold`` (PCC, agreement, IoU, recall); ``"max"``: pass when + ``value <= threshold`` (errors, ADE/FDE). NaN never passes.""" + + threshold: float + direction: str = "min" + metric: str = "pcc" + + def __post_init__(self) -> None: + if self.direction not in ("min", "max"): + raise ValueError(f"gate direction {self.direction!r}: expected 'min' or 'max'") + object.__setattr__(self, "threshold", float(self.threshold)) + + @classmethod + def of(cls, spec: Union["Gate", float, int, Sequence, Mapping], metric: Optional[str] = None) -> "Gate": + """``0.999`` | ``(0.999, "min")`` | ``(0.05, "max", "max_abs")`` | ``{"threshold": ..., ...}`` | Gate. + A bare number takes the direction of ``metric`` (default ``pcc``) from :data:`DIRECTIONS` and raises + ``ValueError`` for a metric whose direction is unknown; a direction that contradicts a known metric's + (e.g. a ``"max"`` PCC gate) raises too.""" + if isinstance(spec, Gate): + gate = spec + elif isinstance(spec, Mapping): + gate = cls(**spec) + elif isinstance(spec, (int, float)): + name = metric or "pcc" + if name not in DIRECTIONS: + raise ValueError(f"gate {spec!r} on metric {name!r}: its direction is unknown; give " + f"(threshold, 'min' | 'max', {name!r})") + gate = cls(float(spec), DIRECTIONS[name], name) + else: + parts = list(spec) + gate = cls(*parts) if len(parts) == 3 else cls(parts[0], parts[1], metric or "pcc") + known = DIRECTIONS.get(gate.metric) + if known is not None and gate.direction != known: + better = "higher" if known == "min" else "lower" + raise ValueError(f"gate {gate}: {gate.metric!r} is a {known!r} metric ({better} is better)") + return gate + + def passes(self, value: float) -> bool: + if value is None or (isinstance(value, float) and math.isnan(value)): + return False + return value >= self.threshold if self.direction == "min" else value <= self.threshold + + def looser_than(self, other: "Gate") -> bool: + """True if this gate accepts something ``other`` rejects (same direction), or the direction differs.""" + if self.direction != other.direction or self.metric != other.metric: + return True + return self.threshold < other.threshold if self.direction == "min" else self.threshold > other.threshold + + +@dataclass +class GateResult: + name: str + value: Optional[float] + gate: Gate + passed: bool + newly_frozen: bool = False + + def __str__(self) -> str: + op = ">=" if self.gate.direction == "min" else "<=" + v = "nan" if self.value is None else f"{self.value:.6g}" + return f"{'PASS' if self.passed else 'FAIL'} {self.name}: {self.gate.metric} {v} {op} {self.gate.threshold:g}" + + +class GateLoosenedError(AssertionError): + """A declared gate is looser than the threshold frozen at the first green run.""" + + +class GateRegistry: + """Declared gates + the JSON record of frozen gates next to the test file. + + Usage in a device test:: + + GATES = GateRegistry.for_test(__file__, {"backbone": 0.999, "head.reg": 0.99, "dets.recall": (0.95, "min")}) + + @pytest.mark.parametrize("stage", GATES.names()) + def test_stage(stage, stage_outputs): + GATES.require(stage, pcc(*stage_outputs[stage])) # AssertionError on failure; frozen on first pass + + ``write=False`` (or ``TTAW_GATES_READONLY=1``) never writes the JSON (CI / verification runs).""" + + VERSION = 1 + + def __init__(self, path: PathLike, gates: Optional[Mapping[str, Any]] = None, *, write: Optional[bool] = None): + self.path = Path(path) + if write is None: + write = os.environ.get("TTAW_GATES_READONLY", "0").strip() in ("", "0", "false", "no") + self.write = bool(write) + self.declared: Dict[str, Gate] = {k: Gate.of(v) for k, v in (gates or {}).items()} + self.frozen: Dict[str, Dict[str, Any]] = {} + if self.path.is_file(): + data = json.loads(self.path.read_text()) + self.frozen = dict(data.get("gates", {})) + self.results: List[GateResult] = [] + loosened = [] + for name, g in self.declared.items(): + if name in self.frozen and g.looser_than(self._frozen_gate(name)): + f = self._frozen_gate(name) + loosened.append(f"{name}: declared {g.direction} {g.threshold:g} ({g.metric}) vs frozen " + f"{f.direction} {f.threshold:g} ({f.metric})") + if loosened: + raise GateLoosenedError(f"gates in {self.path.name} may never be loosened once recorded: " + + "; ".join(loosened)) + + @classmethod + def for_test(cls, test_file: PathLike, gates: Optional[Mapping[str, Any]] = None, **kwargs: Any) -> "GateRegistry": + """Registry stored as ``.gates.json`` next to ``test_file`` (pass ``__file__``).""" + p = Path(test_file) + return cls(p.with_name(p.stem + ".gates.json"), gates, **kwargs) + + def _frozen_gate(self, name: str) -> Gate: + rec = self.frozen[name] + return Gate(rec["threshold"], rec["direction"], rec["metric"]) + + def names(self) -> List[str]: + return sorted(set(self.declared) | set(self.frozen)) + + def gate(self, name: str) -> Gate: + """The effective gate: the declared one (never looser than the frozen one), else the frozen one.""" + if name in self.declared: + return self.declared[name] + if name in self.frozen: + return self._frozen_gate(name) + raise KeyError(f"no gate {name!r} declared or frozen in {self.path.name}") + + def check(self, name: str, value: Optional[float], gate: Optional[Any] = None) -> GateResult: + """Evaluate ``value``; on a pass, freeze the gate if it is new or tighter than the frozen one.""" + g = Gate.of(gate) if gate is not None else self.gate(name) + if name in self.frozen and g.looser_than(self._frozen_gate(name)): + raise GateLoosenedError(f"gate {name!r} {g} is looser than the frozen {self.frozen[name]}") + v = None if value is None else float(value) + passed = g.passes(v) + newly = False + if passed and (name not in self.frozen or self._frozen_gate(name) != g): + self.frozen[name] = {**asdict(g), "value_at_freeze": v, "frozen_at": _now()} + newly = True + if self.write: + self.save() + result = GateResult(name, v, g, passed, newly) + self.results.append(result) + return result + + def require(self, name: str, value: Optional[float], gate: Optional[Any] = None) -> GateResult: + """:meth:`check`, raising ``AssertionError`` when the gate fails.""" + result = self.check(name, value, gate) + if not result.passed: + raise AssertionError(str(result)) + return result + + def save(self) -> None: + body = {"version": self.VERSION, "note": "Frozen at the first green run; never loosen (PLAN.md 0.3).", + "gates": dict(sorted(self.frozen.items()))} + _atomic_write_text(self.path, json.dumps(body, indent=1, sort_keys=False) + "\n") + + +# ----------------------------------------------------------------------------------------------- reports + +@dataclass +class TapComparison: + name: str + metric: str + value: Optional[float] + gate: Optional[Dict[str, Any]] + passed: Optional[bool] + stats: Dict[str, Any] = field(default_factory=dict) + note: str = "" + + +class CompareReport: + """Per-tap comparison records -> JSON (``logs//compare_.json``) and a text table.""" + + def __init__(self, title: str = "", meta: Optional[Mapping[str, Any]] = None, + gates: Optional[GateRegistry] = None): + self.title = title + self.meta = dict(meta or {}) + self.gates = gates + self.entries: List[TapComparison] = [] + + def _gate_for(self, name: str, gate: Any, metric: str) -> Optional[Gate]: + if gate is not None: + return Gate.of(gate, metric) + if self.gates is not None and name in self.gates.names(): + return self.gates.gate(name) + return None + + def add(self, name: str, test: Any, ref: Any, *, metric: str = "pcc", gate: Any = None, + **metric_kwargs: Any) -> TapComparison: + """Compare two tensors with ``metric`` (a :data:`METRICS` name); always records the error statistics.""" + t, r = _host(test), _host(ref) + stats: Dict[str, Any] = {"shape_test": list(t.shape), "shape_ref": list(r.shape), + "dtype_test": str(t.dtype), "dtype_ref": str(r.dtype)} + note = "" + if t.shape == r.shape: + stats.update(_m.error_stats(t, r)) + fn = METRICS[metric][0] + value: Optional[float] = float(fn(t, r, **metric_kwargs)) + else: + value, note = None, "shape mismatch" + return self._record(name, metric, value, gate, stats, note) + + def add_value(self, name: str, value: Optional[float], *, metric: str, gate: Any = None, + stats: Optional[Mapping[str, Any]] = None) -> TapComparison: + """Record an already computed scalar (detection recall, ADE, argmax agreement ...).""" + return self._record(name, metric, None if value is None else float(value), gate, dict(stats or {}), "") + + def _record(self, name: str, metric: str, value: Optional[float], gate: Any, stats: Dict[str, Any], + note: str) -> TapComparison: + g = self._gate_for(name, gate, metric) + passed: Optional[bool] = None + if g is not None: + passed = (self.gates.check(name, value, g).passed if self.gates is not None else g.passes(value)) + entry = TapComparison(name, metric, value, None if g is None else asdict(g), passed, stats, note) + self.entries.append(entry) + return entry + + @property + def passed(self) -> bool: + """All gated entries passed (ungated entries are informational).""" + return all(e.passed for e in self.entries if e.passed is not None) + + def to_dict(self) -> Dict[str, Any]: + return {"title": self.title, "created": _now(), "meta": self.meta, "passed": self.passed, + "entries": [asdict(e) for e in self.entries]} + + def save_json(self, path: PathLike) -> Path: + path = Path(path) + _atomic_write_text(path, json.dumps(self.to_dict(), indent=1, default=str) + "\n") + return path + + def table(self) -> str: + rows = [f"{self.title}".strip(), f"{'name':40s} {'metric':16s} {'value':>12s} {'gate':>14s} verdict"] + for e in self.entries: + gate = "" + if e.gate is not None: + gate = (">=" if e.gate["direction"] == "min" else "<=") + f"{e.gate['threshold']:g}" + val = "nan" if e.value is None else f"{e.value:.6g}" + verdict = {True: "PASS", False: "FAIL", None: "-"}[e.passed] + rows.append(f"{e.name:40s} {e.metric:16s} {val:>12s} {gate:>14s} {verdict} {e.note}".rstrip()) + return "\n".join(r for r in rows if r) + + +def compare_taps(test: Mapping[str, Any], golden: Mapping[str, Any], *, gates: Optional[Any] = None, + metric: str = "pcc", names: Optional[Sequence[str]] = None, title: str = "", + meta: Optional[Mapping[str, Any]] = None) -> CompareReport: + """Compare every tap present in both mappings (or ``names``). ``gates``: a :class:`GateRegistry` or a + ``{name: gate spec}`` mapping. Taps missing on either side are recorded as failures when gated.""" + registry = gates if isinstance(gates, GateRegistry) else None + plain = {} if registry is not None or gates is None else {k: Gate.of(v, metric) for k, v in gates.items()} + report = CompareReport(title, meta, registry) + for name in (names if names is not None else [n for n in golden if n in test]): + if name not in test or name not in golden: + report.add_value(name, None, metric=metric, gate=plain.get(name), + stats={"missing": "test" if name not in test else "golden"}) + continue + report.add(name, test[name], golden[name], metric=metric, gate=plain.get(name)) + return report diff --git a/code/tt_diffusion_planner/ttaw/image.py b/code/tt_diffusion_planner/ttaw/image.py new file mode 100644 index 0000000000000000000000000000000000000000..438db2933fdb08d3d61851c7a252b84c37093566 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/image.py @@ -0,0 +1,666 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C14 image pre-processing: bit-exact host emulations of the Autoware camera nodes' resize kernels, driven by +per-source-size lookup tables (LUTs) that are computed once and cached. + +Each preset reproduces one deployed CUDA / OpenCV routine exactly (integer outputs equal, not "close"), because a +standard resize in its place changes the network input enough to fail the PCC gates (YOLOX: ``cv2.resize`` drops the +detection-output PCC to 0.94, research/yolox/SPEC.md section 3). The tables hold everything that depends only on the +source and destination sizes (source indices, interpolation weights, letterbox geometry); applying them is a few +vectorised numpy passes over the image. numpy only; no ttnn, no torch, no OpenCV. + +Presets (PLAN.md C14; the other camera ports add theirs here): + +========================== ========================================================================================= +``yolox_letterbox`` autoware_tensorrt_yolox ``resize_bilinear_letterbox_nhwc_to_nchw32_batch_kernel`` + (autoware_universe @ 9ceaccf, ``perception/autoware_tensorrt_yolox/src/preprocess.cu:43-129``, + launched by ``tensorrt_yolox.cpp:243-332``): a "bilinear" resize with inverted, + half-magnitude weights, float -> int truncation between the horizontal and vertical passes, + ``lroundf``, top-left letterbox padded with 114, BGR kept, no normalisation. +``bevdet_nearest_crop`` the ``bevdet::Preprocess`` TensorRT plugin of autoware_tensorrt_bevdet (``bevdet_vendor`` + 0.2.1; research/bevdet/SPEC.md section 3.2): nearest resize by r = 0.44f + crop (140, 0) + of the 1600x900 image, ``roundf(i / r + crop_h / r)`` rows / columns in float32, planar + B, G, R uint8 (the normalisation ``bevdet_normalize`` is separate: device-side in the TT + port). The node's ``cv::resize`` of non-1600x900 images to 1600x900 stays with the caller. +``bevformer_preprocess`` autoware_tensorrt_bevformer's ``preprocessing_pipeline`` (autoware_universe @ 9ceaccf, + ``src/preprocessing/normalize_multiview_image.cpp:70-95`` + + ``augmentation_transforms.cpp:90-222``; research/bevformer/SPEC.md section 3.1): + normalise FIRST (BGR uint8 -> float, ``(x - mean_c) / std_c``, BGR mean, std 1), then + ``cv::resize`` x0.8 with INTER_LINEAR **on float data** (1600x900 -> 1280x720), then pad + bottom / right with 0 (normalised space) to a multiple of 32 (1280x736), CHW. The node's + ``cv::resize`` of non-1600x900 uint8 images to 1600x900 stays with the caller. +========================== ========================================================================================= + +Exactness notes for ``yolox_letterbox`` (all arithmetic is float32 in the kernel): + +- ``scale = min(dst_w / (float)src_w, dst_h / (float)src_h)``; ``r_h = (int)(scale * src_h)``, + ``r_w = (int)(scale * src_w)``; the kernel receives ``k = (float)(1.0 / (double)scale)`` (``preprocess.cu:121-128``). +- per output row ``h``: ``c = k * (float)(h + 0.5)``, ``src_row = clamp(lroundf(c) - 1, 0, src_h - 2)``, + weight ``t = (1 + lroundf(c) - c) / 2`` (float32), and the same per column (``preprocess.cu:48-59, 77-88``). +- ``lerp1d(a, b, t) = fma(t, b, fma(-t, a, a))`` with float operands (CUDA's float overload, i.e. ``fmaf``: + one rounding per call); the horizontal results are truncated to int before the vertical ``lerp1d`` + (``lerp1d`` takes ``int a, int b``), and the result is ``lroundf``-ed (``preprocess.cu:43-46, 57, 101-103``). +- ``t`` lies in [0.25, 0.75], so ``t`` is a multiple of 2**-25 and ``a - t*a`` / ``t*b + inner`` with integer + ``a, b`` in [0, 255] need at most 34 significant bits: float64 holds them exactly, and one ``astype(float32)`` is + exactly the single rounding of ``fmaf``. +- Rows ``h >= r_h`` and columns ``w >= r_w`` are 114 (``preprocess.cu:109-110``); the output values are integers in + [0, 255], so this module returns uint8 (the network input is ``float(value)``, ``norm_factor`` 1.0). + +Verified against two independent scalar transliterations of the CUDA source (``common/tests/host/test_image_host.py`` +and research/yolox/scripts/{check_preprocess,verify_preproc_scalar}.py: 0 mismatches). Residual uncertainty, from the +SPEC: if nvcc picked the double ``fma`` overload, 0-2 of 3000 samples would differ by 1 (not testable without CUDA). + +Exactness notes for ``bevformer_preprocess`` (float32 data; OpenCV's float INTER_LINEAR path): + +- normalisation: ``Mat - mean`` / ``std`` evaluates as ``convertTo(alpha = 1/std, beta = -mean/std)`` with float + ``alpha`` / ``beta`` (one fused multiply-add per value); with the deployed std 1 it is ``x - float(mean_c)``. +- resize coefficients (``resize.cpp``): ``scale = 1.0 / ((double)dst / src)``, ``f = (float)((d + 0.5) * scale - + 0.5)``, ``s = floor(f)``, ``w = f - s`` (float), clamped to the border (``s < 0`` -> ``s = 0, w = 0``; + ``s >= src - 1`` -> ``s = src - 1, w = 0``); the destination size is ``(int)(src * 0.8f)`` in float. +- each pass is ``fma(x1 - x0, w, x0)`` in float32 (the difference rounded once, then one rounding of the fused + product-sum), horizontal pass first, then vertical. This is what OpenCV 4.8.1 and 4.11.0 compute on this x86-64 + host (AVX2 + FMA) bit for bit for every x0.8 down-scale, on arbitrary float images (not just uint8 - mean): the + plain two-product form ``x0 * (1 - w) + x1 * w`` differs from it in about 10 % of the samples by 1 ulp + (``tests/host/test_image_bevformer_host.py`` holds both checks). The weights of a x0.8 resize are multiples of + 1/8, so float64 holds every fused product-sum exactly and one ``astype(float32)`` is the single rounding. OpenCV + computes the coefficients of other scales differently, so :func:`opencv_linear_lut` refuses them by default. +""" +from __future__ import annotations + +import threading +from collections import OrderedDict +from dataclasses import dataclass +from typing import Any, Callable, Dict, Hashable, Optional, Tuple + +import numpy as np + +__all__ = [ + "LetterboxGeometry", + "YoloxLetterboxLUT", + "yolox_letterbox", + "yolox_letterbox_geometry", + "yolox_letterbox_lut", + "lroundf", + "f32", + "LUTCache", + "PRESETS", + "YOLOX_PAD_VALUE", + "NearestCropLUT", + "bevdet_nearest_lut", + "bevdet_nearest_crop", + "bevdet_normalize", + "BEVDET_MEAN", + "BEVDET_STD", + "LinearResizeLUT", + "opencv_linear_lut", + "opencv_resize_linear_f32", + "bevformer_input_geometry", + "bevformer_normalize", + "bevformer_preprocess", + "BEVFORMER_MEAN_BGR", + "BEVFORMER_STD", +] + +F32 = np.float32 +YOLOX_PAD_VALUE = 114 # preprocess.cu:109-110 (letterbox fill), the YOLOX training pad value + + +def f32(x: Any) -> np.ndarray: + """``x`` rounded once to float32 (round-to-nearest-even), as a numpy float32 scalar or array.""" + return np.asarray(x).astype(F32) + + +def lroundf(x: Any) -> np.ndarray: + """C ``lroundf``: round half away from zero, for float32 inputs (exact in float64). Returns int64.""" + x = np.asarray(x, dtype=np.float64) + return (np.sign(x) * np.floor(np.abs(x) + 0.5)).astype(np.int64) + + +class LUTCache: + """A small thread-safe LRU of tables keyed by the source / destination sizes (a camera stream keeps its size, + so the table is built once per stream; Autoware re-creates its buffers only when the source size changes, + ``tensorrt_yolox.cpp:253-291``).""" + + def __init__(self, maxsize: int = 8): + self.maxsize = int(maxsize) + self._items: "OrderedDict[Hashable, Any]" = OrderedDict() + self._lock = threading.Lock() + self.hits = 0 + self.misses = 0 + + def get(self, key: Hashable, make: Callable[[], Any]) -> Any: + with self._lock: + if key in self._items: + self._items.move_to_end(key) + self.hits += 1 + return self._items[key] + value = make() # outside the lock: building a table takes a few ms + with self._lock: + self.misses += 1 + self._items[key] = value + self._items.move_to_end(key) + while len(self._items) > self.maxsize: + self._items.popitem(last=False) + return value + + def clear(self) -> None: + with self._lock: + self._items.clear() + self.hits = self.misses = 0 + + def __len__(self) -> int: + return len(self._items) + + +# ------------------------------------------------------------------------------------------- YOLOX letterbox + + +@dataclass(frozen=True) +class LetterboxGeometry: + """Where the source image lands in the network input, and what the decoder and the mask crop use. + + ``scale`` is the float32 scale of ``preprocess.cu:121`` (also ``scales_`` of ``tensorrt_yolox.cpp:285``, used to + map boxes back, and the mask scale of ``tensorrt_yolox.cpp:528-536``); ``resized_hw`` = ``(r_h, r_w)``, the + un-letterboxed region at the top-left of the ``dst_hw`` canvas. Autoware's mask crop ``(out_h, out_w)`` = + ``((int)(src_h * scale), (int)(src_w * scale))`` is the same pair (float multiplication commutes).""" + + src_hw: Tuple[int, int] + dst_hw: Tuple[int, int] + scale: float # exact float32 value, stored as a Python float + resized_hw: Tuple[int, int] + pad_value: int = YOLOX_PAD_VALUE + + @property + def scale_f32(self) -> np.float32: + return F32(self.scale) + + @property + def mask_hw(self) -> Tuple[int, int]: + """``(out_h, out_w)`` of Autoware's mask (``getMaskImageGpu``): the un-letterboxed region.""" + return self.resized_hw + + def to_dict(self) -> Dict[str, Any]: + return {"src_hw": list(self.src_hw), "dst_hw": list(self.dst_hw), "scale": self.scale, + "scale_f32_hex": F32(self.scale).tobytes()[::-1].hex(), "resized_hw": list(self.resized_hw), + "mask_hw": list(self.mask_hw), "pad_value": self.pad_value, + "letterbox": "top-left (padding at the bottom / right)"} + + +def yolox_letterbox_geometry(src_h: int, src_w: int, dst_h: int = 960, dst_w: int = 960, + pad_value: int = YOLOX_PAD_VALUE) -> LetterboxGeometry: + """Letterbox geometry of ``preprocess.cu:121-123`` (float32 scale, int truncation).""" + src_h, src_w, dst_h, dst_w = int(src_h), int(src_w), int(dst_h), int(dst_w) + if src_h < 2 or src_w < 2: + raise ValueError(f"the Autoware resize needs a source of at least 2x2 pixels, got {src_w}x{src_h}") + if dst_h < 1 or dst_w < 1: + raise ValueError(f"bad destination size {dst_w}x{dst_h}") + scale = min(F32(dst_w) / F32(src_w), F32(dst_h) / F32(src_h)) # float32 divisions + scale = F32(scale) + r_h = int(F32(scale * F32(src_h))) + r_w = int(F32(scale * F32(src_w))) + return LetterboxGeometry((src_h, src_w), (dst_h, dst_w), float(scale), (min(r_h, dst_h), min(r_w, dst_w)), + int(pad_value)) + + +def _axis_table(n_dst: int, n_src: int, k: np.float32) -> Tuple[np.ndarray, np.ndarray]: + """Source index and weight of every destination row (or column): ``preprocess.cu:48-51, 77-88``.""" + c = (k * (np.arange(n_dst) + 0.5).astype(F32)).astype(F32) # centroid = scale * (float)(h + 0.5) + rc = lroundf(c) + idx = rc - 1 + idx = np.where(idx < 0, 0, idx) + idx = np.where(idx >= n_src - 1, n_src - 2, idx) + weight = ((F32(1) + rc.astype(F32)).astype(F32) - c).astype(F32) # 1 + lroundf(c) - c (float32) + weight = (weight / F32(2)).astype(F32) + return idx.astype(np.int64), weight + + +@dataclass(frozen=True) +class YoloxLetterboxLUT: + """The per-size tables of the YOLOX letterbox: source row / column of every destination row / column and the + float32 weights. Build with :func:`yolox_letterbox_lut` (cached); apply with :meth:`apply`.""" + + geometry: LetterboxGeometry + rows: np.ndarray # (r_h,) int64: source row hi (hi + 1 is the second tap) + row_w: np.ndarray # (r_h,) float32: vertical weight + cols: np.ndarray # (r_w,) int64: source column wi + col_w: np.ndarray # (r_w,) float32: horizontal weight + k: float # the kernel's float32 ``scale`` argument (1 / scale) + + def apply(self, image: np.ndarray, *, out: Optional[np.ndarray] = None, chunk_rows: int = 96) -> np.ndarray: + """Letterbox one BGR uint8 image ``(src_h, src_w, 3)`` (any channel count) -> uint8 ``(dst_h, dst_w, C)`` + HWC, channel order unchanged. ``out`` (uint8, that shape) is written in place when given.""" + g = self.geometry + img = np.asarray(image) + if img.ndim != 3 or img.dtype != np.uint8 or tuple(img.shape[:2]) != g.src_hw: + raise ValueError(f"expected a uint8 ({g.src_hw[0]}, {g.src_hw[1]}, C) image, got {img.dtype} {img.shape}") + dst_h, dst_w = g.dst_hw + r_h, r_w = g.resized_hw + shape = (dst_h, dst_w, img.shape[2]) + if out is None: + out = np.empty(shape, np.uint8) + elif out.shape != shape or out.dtype != np.uint8: + raise ValueError(f"out must be uint8 {shape}, got {out.dtype} {out.shape}") + out[r_h:] = g.pad_value + out[:r_h, r_w:] = g.pad_value + if r_h == 0 or r_w == 0: + return out + tw = self.col_w.astype(np.float64)[None, :, None] # exact float32 values + cols0, cols1 = self.cols, self.cols + 1 + for start in range(0, r_h, max(1, int(chunk_rows))): + stop = min(r_h, start + int(chunk_rows)) + hi = self.rows[start:stop] + # horizontal pass on the source rows this chunk needs (hi and hi + 1), each row once + need = np.unique(np.concatenate([hi, hi + 1])) + src = img[need] # (R, src_w, C) uint8 + a = src[:, cols0].astype(np.float64) # (R, r_w, C) + b = src[:, cols1].astype(np.float64) + inner = (a - tw * a).astype(F32) # fmaf(-t, a, a) + horiz = (tw * b + inner).astype(F32) # fmaf(t, b, inner) + horiz = np.trunc(horiz).astype(np.float64) # lerp1d(int a, int b, ...): truncation + pos0 = np.searchsorted(need, hi) + pos1 = np.searchsorted(need, hi + 1) + a1, b1 = horiz[pos0], horiz[pos1] # (rows, r_w, C) + th = self.row_w[start:stop].astype(np.float64)[:, None, None] + inner = (a1 - th * a1).astype(F32) + r = (th * b1 + inner).astype(F32) # second lerp1d, float32 + out[start:stop, :r_w] = np.floor(r.astype(np.float64) + 0.5) # lroundf (r >= 0) + return out + + +_YOLOX_LUTS = LUTCache(maxsize=8) + + +def yolox_letterbox_lut(src_h: int, src_w: int, dst_h: int = 960, dst_w: int = 960, + pad_value: int = YOLOX_PAD_VALUE) -> YoloxLetterboxLUT: + """The (cached) tables for one source size.""" + + def make() -> YoloxLetterboxLUT: + g = yolox_letterbox_geometry(src_h, src_w, dst_h, dst_w, pad_value) + k = F32(1.0 / np.float64(g.scale_f32)) # 1.0 / scale in double, passed as a float argument + rows, row_w = _axis_table(dst_h, g.src_hw[0], k) + cols, col_w = _axis_table(dst_w, g.src_hw[1], k) + r_h, r_w = g.resized_hw + return YoloxLetterboxLUT(g, rows[:r_h].copy(), row_w[:r_h].copy(), cols[:r_w].copy(), col_w[:r_w].copy(), + float(k)) + + return _YOLOX_LUTS.get(("yolox", int(src_h), int(src_w), int(dst_h), int(dst_w), int(pad_value)), make) + + +def yolox_letterbox(image: np.ndarray, dst_hw: Tuple[int, int] = (960, 960), *, pad_value: int = YOLOX_PAD_VALUE, + layout: str = "hwc", out: Optional[np.ndarray] = None) -> Tuple[np.ndarray, LetterboxGeometry]: + """Autoware YOLOX letterbox of one BGR8 image ``(H, W, 3)`` -> ``(uint8 image, geometry)``. + + ``layout``: ``"hwc"`` -> ``(dst_h, dst_w, 3)``; ``"nchw"`` -> ``(1, 3, dst_h, dst_w)`` (the ONNX ``images`` + layout; ``.astype(np.float32)`` is exactly the network input). The channel order is unchanged (Autoware feeds + BGR).""" + img = np.asarray(image) + lut = yolox_letterbox_lut(img.shape[0], img.shape[1], int(dst_hw[0]), int(dst_hw[1]), pad_value) + if layout == "hwc": + return lut.apply(img, out=out), lut.geometry + if layout == "nchw": + hwc = lut.apply(img) + res = np.ascontiguousarray(hwc.transpose(2, 0, 1)[None]) + if out is not None: + out[...] = res + return out, lut.geometry + return res, lut.geometry + raise ValueError(f"layout must be 'hwc' or 'nchw', got {layout!r}") + + +# ------------------------------------------------------------------------------------------- BEVDet nearest crop + +# Per-plane statistics of the BEVDet Preprocess plugin, applied to the B, G, R planes in this order: training loaded +# images as RGB (PIL) and its mmlabNormalize swapped them to BGR before applying the "RGB" ImageNet statistics, and +# the Autoware node feeds BGR planes with the same numbers (research/bevdet/SPEC.md section 3.1). +BEVDET_MEAN = (123.675, 116.28, 103.53) +BEVDET_STD = (58.395, 57.12, 57.375) + + +def _roundf(x: Any) -> np.ndarray: + """C ``roundf`` (half away from zero) of float32 values -> int64.""" + x = np.asarray(x, F32) + return (np.sign(x) * np.floor(np.abs(x) + F32(0.5))).astype(np.int64) + + +@dataclass(frozen=True) +class NearestCropLUT: + """The ``bevdet::Preprocess`` gather: output row ``i`` / column ``j`` reads source row ``rows[i]`` / column + ``cols[j]`` of the (already 1600x900) image. Build with :func:`bevdet_nearest_lut` (cached); apply with + :meth:`apply` (one HWC image) or :meth:`apply_planes` (planar ``[..., C, H, W]``).""" + + src_hw: Tuple[int, int] + dst_hw: Tuple[int, int] + crop_hw: Tuple[int, int] + resize: float # exact float32 ratio dst_w / src_w, stored as a Python float + rows: np.ndarray # (dst_h,) int64 + cols: np.ndarray # (dst_w,) int64 + + def apply(self, image: np.ndarray, *, channels: str = "rgb", out: Optional[np.ndarray] = None) -> np.ndarray: + """One ``(src_h, src_w, 3)`` uint8 image -> ``(3, dst_h, dst_w)`` uint8 B, G, R planes. ``channels`` is the + input's order: ``"rgb"`` (PIL / ``ttaw.io.load_image``) or ``"bgr"`` (OpenCV, cv_bridge ``bgr8``).""" + img = np.asarray(image) + if img.ndim != 3 or img.shape[2] != 3 or img.dtype != np.uint8 or tuple(img.shape[:2]) != self.src_hw: + raise ValueError(f"expected a uint8 ({self.src_hw[0]}, {self.src_hw[1]}, 3) image, got {img.dtype} " + f"{img.shape} (resize it to the source size first, like the node's cv::resize)") + if channels not in ("rgb", "bgr"): + raise ValueError(f"channels must be 'rgb' or 'bgr', got {channels!r}") + sub = img[self.rows][:, self.cols] # (dst_h, dst_w, 3) + order = [2, 1, 0] if channels == "rgb" else [0, 1, 2] # -> B, G, R + res = sub[:, :, order].transpose(2, 0, 1) + if out is None: + return np.ascontiguousarray(res) + shape = (3,) + tuple(self.dst_hw) + if out.shape != shape or out.dtype != np.uint8: + raise ValueError(f"out must be uint8 {shape}, got {out.dtype} {out.shape}") + out[...] = res + return out + + def apply_planes(self, planes: np.ndarray) -> np.ndarray: + """Planar ``[..., C, src_h, src_w]`` (any dtype) -> ``[..., C, dst_h, dst_w]``.""" + p = np.asarray(planes) + if tuple(p.shape[-2:]) != self.src_hw: + raise ValueError(f"expected planes [..., {self.src_hw[0]}, {self.src_hw[1]}], got {p.shape}") + return np.ascontiguousarray(p[..., self.rows, :][..., self.cols]) + + +_BEVDET_LUTS = LUTCache(maxsize=4) + + +def bevdet_nearest_lut(src_hw: Tuple[int, int] = (900, 1600), dst_hw: Tuple[int, int] = (256, 704), + crop_hw: Tuple[int, int] = (140, 0)) -> NearestCropLUT: + """The (cached) gather of the BEVDet Preprocess plugin: ``r = (float)dst_w / src_w`` (0.44f for the deployed + config, also the ONNX attribute ``resize_radio``), ``rows[i] = roundf(i / r + crop_h / r)``, + ``cols[j] = roundf(j / r + crop_w / r)``, every quantity float32 (deployed: rows 318, 320, 323, ..., 898 and + columns 0, 2, 5, 7, ..., 1598). ``crop_hw`` is the crop offset in RESIZED pixels (the yaml ``crop``, (h, w)).""" + src_h, src_w = int(src_hw[0]), int(src_hw[1]) + dst_h, dst_w = int(dst_hw[0]), int(dst_hw[1]) + crop_h, crop_w = int(crop_hw[0]), int(crop_hw[1]) + + def make() -> NearestCropLUT: + r = F32(F32(dst_w) / F32(src_w)) + off_h = F32(F32(crop_h) / r) + off_w = F32(F32(crop_w) / r) + rows = _roundf((np.arange(dst_h, dtype=F32) / r + off_h).astype(F32)) + cols = _roundf((np.arange(dst_w, dtype=F32) / r + off_w).astype(F32)) + if rows.min() < 0 or rows.max() >= src_h or cols.min() < 0 or cols.max() >= src_w: + raise ValueError(f"crop {crop_hw} with ratio {float(r)} reads outside the {src_w}x{src_h} source") + return NearestCropLUT((src_h, src_w), (dst_h, dst_w), (crop_h, crop_w), float(r), rows, cols) + + return _BEVDET_LUTS.get(("bevdet", src_h, src_w, dst_h, dst_w, crop_h, crop_w), make) + + +def bevdet_nearest_crop(image: np.ndarray, *, channels: str = "rgb", src_hw: Tuple[int, int] = (900, 1600), + dst_hw: Tuple[int, int] = (256, 704), crop_hw: Tuple[int, int] = (140, 0), + out: Optional[np.ndarray] = None) -> np.ndarray: + """One camera image at the source size (1600x900) -> its ``(3, 256, 704)`` uint8 B, G, R crop: exactly the + samples the BEVDet network sees before ``bevdet_normalize``.""" + return bevdet_nearest_lut(src_hw, dst_hw, crop_hw).apply(image, channels=channels, out=out) + + +def bevdet_normalize(crop: np.ndarray, mean: Any = BEVDET_MEAN, std: Any = BEVDET_STD) -> np.ndarray: + """``(x - mean[c]) / std[c]`` in float32 on B, G, R planes ``[..., 3, H, W]`` (the plugin's fp32 path; its fp16 + mode computes in ``__half``).""" + m = np.asarray(mean, F32).reshape(3, 1, 1) + s = np.asarray(std, F32).reshape(3, 1, 1) + return ((np.asarray(crop).astype(F32) - m) / s).astype(F32) + + +# ------------------------------------------------------------------------------------------- BEVFormer + +# autoware_tensorrt_bevformer config/bevformer.param.yaml data_params (= the BEVFormer training img_norm_cfg, +# DerryHub configs/bevformer/bevformer_small.py:19): caffe-style BGR means, std 1, to_rgb false +BEVFORMER_MEAN_BGR = (103.530, 116.280, 123.675) +BEVFORMER_STD = (1.0, 1.0, 1.0) + + +@dataclass(frozen=True) +class LinearResizeLUT: + """OpenCV ``cv::resize(..., INTER_LINEAR)`` of float32 data for one (source, destination) size pair: row / column + pairs and the float32 weight of the second tap (``resize.cpp`` coefficient loop; module docstring). Build with + :func:`opencv_linear_lut` (cached); apply with :meth:`apply` (HWC) or :meth:`apply_planes` (``[..., H, W]``).""" + + src_hw: Tuple[int, int] + dst_hw: Tuple[int, int] + rows0: np.ndarray # (dst_h,) int64 + rows1: np.ndarray # (dst_h,) int64 (clamped to the last row) + wy: np.ndarray # (dst_h,) float32 weight of rows1 + cols0: np.ndarray # (dst_w,) int64 + cols1: np.ndarray + wx: np.ndarray # (dst_w,) float32 weight of cols1 + + @staticmethod + def _lerp(x0: np.ndarray, x1: np.ndarray, w: Any, out: Optional[np.ndarray] = None) -> np.ndarray: + """``fma(x1 - x0, w, x0)`` in float32: the difference rounded once, the fused product-sum rounded once + (float64 holds ``d * w + x0`` exactly for weights that are multiples of 1/8). ``x1`` is overwritten.""" + d = np.subtract(x1, x0, out=x1) # float32: one rounding + t = d.astype(np.float64) + t *= w + t += x0 + if out is None: + return t.astype(F32) + out[...] = t # float64 -> float32: one rounding + return out + + @staticmethod + def _period(i0: np.ndarray, i1: np.ndarray, w: np.ndarray, n_src: int) -> Optional[Tuple[int, int]]: + """``(p_in, p_out)`` when the taps repeat every ``p_out`` outputs shifted by ``p_in`` inputs, both taps of + an output lie in one input block and no border clamp occurs (x0.8: 5 inputs -> 4 outputs), else None.""" + n_dst = len(i0) + for p_out in range(1, min(n_dst, 64) + 1): + if n_dst % p_out or (n_src * p_out) % n_dst: + continue + p_in = n_src * p_out // n_dst + base = (np.arange(n_dst) // p_out) * p_in + reps = n_dst // p_out + if (np.array_equal(i0 - base, np.tile(i0[:p_out], reps)) and np.array_equal(i1, i0 + 1) + and np.array_equal(w, np.tile(w[:p_out], reps)) and int(i1[:p_out].max()) < p_in): + return p_in, p_out + return None + + def _pass(self, p: np.ndarray, axis: int, i0: np.ndarray, i1: np.ndarray, w: np.ndarray) -> np.ndarray: + """One interpolation pass along ``axis`` (-1: columns, -2: rows) of a contiguous ``p`` [..., H, W]: strided + block slices when the taps are periodic (no transposes), else gathers.""" + period = self._period(i0, i1, w, p.shape[axis]) + if period is None: + wb = w.astype(np.float64) if axis == -1 else w.astype(np.float64)[:, None] + return self._lerp(np.take(p, i0, axis=axis), np.take(p, i1, axis=axis), wb) + p_in, p_out = period + nb = p.shape[axis] // p_in + taps0 = [int(t) for t in i0[:p_out]] + run = taps0 == list(range(taps0[0], taps0[0] + p_out)) # x0.8: taps (0..3) and (1..4) of 5 + wj = np.asarray(w[:p_out], np.float64) + if axis == -1: + blk = p.reshape(p.shape[:-1] + (nb, p_in)) # [..., H, nb, p_in] (a view) + if run: + a = taps0[0] + res = self._lerp(blk[..., a:a + p_out], np.array(blk[..., a + 1:a + 1 + p_out]), wj) + else: + res = np.empty(p.shape[:-1] + (nb, p_out), F32) + for j in range(p_out): + self._lerp(blk[..., taps0[j]], np.array(blk[..., taps0[j] + 1]), wj[j], out=res[..., j]) + return res.reshape(p.shape[:-1] + (nb * p_out,)) + blk = p.reshape(p.shape[:-2] + (nb, p_in, p.shape[-1])) # [..., nb, p_in, W] (a view) + if run: + a = taps0[0] + res = self._lerp(blk[..., a:a + p_out, :], np.array(blk[..., a + 1:a + 1 + p_out, :]), wj[:, None]) + else: + res = np.empty(p.shape[:-2] + (nb, p_out, p.shape[-1]), F32) + for j in range(p_out): + self._lerp(blk[..., taps0[j], :], np.array(blk[..., taps0[j] + 1, :]), wj[j], out=res[..., j, :]) + return res.reshape(p.shape[:-2] + (nb * p_out, p.shape[-1])) + + def apply_planes(self, planes: np.ndarray, *, out: Optional[np.ndarray] = None) -> np.ndarray: + """Planar float32 ``[..., src_h, src_w]`` -> ``[..., dst_h, dst_w]`` (horizontal pass first, as OpenCV).""" + p = np.ascontiguousarray(planes) + if p.dtype != F32 or tuple(p.shape[-2:]) != self.src_hw: + raise ValueError(f"expected float32 planes [..., {self.src_hw[0]}, {self.src_hw[1]}], got {p.dtype} " + f"{p.shape}") + shape = p.shape[:-2] + tuple(self.dst_hw) + if out is not None and (out.shape != shape or out.dtype != F32): + raise ValueError(f"out must be float32 {shape}, got {out.dtype} {out.shape}") + h = self._pass(p, -1, self.cols0, self.cols1, self.wx) # [..., src_h, dst_w] + v = self._pass(h, -2, self.rows0, self.rows1, self.wy) # [..., dst_h, dst_w] + if out is None: + return v + out[...] = v + return out + + def apply(self, image: np.ndarray, *, out: Optional[np.ndarray] = None) -> np.ndarray: + """``(src_h, src_w[, C])`` float32 -> ``(dst_h, dst_w[, C])`` float32 (horizontal pass first, as OpenCV).""" + img = np.asarray(image) + if img.dtype != F32 or tuple(img.shape[:2]) != self.src_hw: + raise ValueError(f"expected a float32 {self.src_hw} image, got {img.dtype} {img.shape}") + planes = np.moveaxis(img, (0, 1), (-2, -1)) # [C..., H, W] + res = np.moveaxis(self.apply_planes(planes), (-2, -1), (0, 1)) + if out is None: + return np.ascontiguousarray(res) + if out.shape != res.shape or out.dtype != F32: + raise ValueError(f"out must be float32 {res.shape}, got {out.dtype} {out.shape}") + out[...] = res + return out + + +def _linear_axis(n_src: int, n_dst: int) -> Tuple[np.ndarray, np.ndarray, np.ndarray]: + """``resize.cpp`` INTER_LINEAR coefficients of one axis: (first tap, second tap, float32 weight of the second).""" + scale = 1.0 / (float(n_dst) / float(n_src)) # inv_scale = (double)dst / src; 1. / inv_scale + f = ((np.arange(n_dst, dtype=np.float64) + 0.5) * scale - 0.5).astype(F32) + s = np.floor(f).astype(np.int64) + w = (f - s.astype(F32)).astype(F32) + low = s < 0 + s[low], w[low] = 0, F32(0) + high = s >= n_src - 1 + s[high], w[high] = n_src - 1, F32(0) + return s, np.minimum(s + 1, n_src - 1), w + + +def _eighths(w: np.ndarray) -> bool: + x = np.asarray(w, np.float64) * 8.0 + return bool(np.all(x == np.round(x))) + + +_LINEAR_LUTS = LUTCache(maxsize=8) + + +def opencv_linear_lut(src_hw: Tuple[int, int], dst_hw: Tuple[int, int], *, strict: bool = True) -> LinearResizeLUT: + """The (cached) float INTER_LINEAR tables for one size pair. + + Verified bit-exact against OpenCV 4.8.1 / 4.11.0 for every x0.8 down-scale (1600x900 -> 1280x720, the deployed + one; 1920x1080, 160x90, 80x45), whose weights are multiples of 1/8. OpenCV computes the coefficients of other + scales differently (up-scales and non-dyadic down-scales differ by up to ~1e-4 relative), so ``strict=True`` + refuses a geometry whose weights are not multiples of 1/8; ``strict=False`` gives the textbook coefficients for + those (an approximation, not an emulation). Exact 2x down-scales are always refused: OpenCV switches + INTER_LINEAR to INTER_AREA there (``resize.cpp``).""" + src_h, src_w = int(src_hw[0]), int(src_hw[1]) + dst_h, dst_w = int(dst_hw[0]), int(dst_hw[1]) + if min(src_h, src_w, dst_h, dst_w) < 1: + raise ValueError(f"empty resize {src_hw} -> {dst_hw}") + if src_h == 2 * dst_h and src_w == 2 * dst_w: + raise ValueError("an exact 2x down-scale runs OpenCV's INTER_AREA path, which is not emulated") + + def make() -> LinearResizeLUT: + r0, r1, wy = _linear_axis(src_h, dst_h) + c0, c1, wx = _linear_axis(src_w, dst_w) + return LinearResizeLUT((src_h, src_w), (dst_h, dst_w), r0, r1, wy, c0, c1, wx) + + lut = _LINEAR_LUTS.get(("linear_f32", src_h, src_w, dst_h, dst_w), make) + if strict and not (_eighths(lut.wx) and _eighths(lut.wy)): + raise ValueError(f"{src_hw} -> {dst_hw} is not a geometry this emulation reproduces bit for bit (weights must " + "be multiples of 1/8, e.g. a x0.8 down-scale); pass strict=False for an approximation") + return lut + + +def opencv_resize_linear_f32(image: np.ndarray, dst_hw: Tuple[int, int], *, strict: bool = True) -> np.ndarray: + """``cv2.resize(image, (dst_w, dst_h))`` (INTER_LINEAR) of a float32 ``(H, W[, C])`` image: bit-exact for the + geometries :func:`opencv_linear_lut` accepts with ``strict=True``.""" + img = np.asarray(image) + return opencv_linear_lut(img.shape[:2], dst_hw, strict=strict).apply(img) + + +def bevformer_input_geometry(src_hw: Tuple[int, int] = (900, 1600), scale: float = 0.8, + pad_divisor: int = 32) -> Dict[str, Tuple[int, int]]: + """Sizes of the node's pipeline for one source size: ``resized_hw`` = ``(int)(src * scale)`` in float32 + (``augmentation_transforms.cpp:104-105``) and ``padded_hw`` (``PadMultiViewImages``, size divisor). Deployed: + (900, 1600) -> (720, 1280) -> (736, 1280); the network normalises its UV by the padded size.""" + s = F32(scale) + h = int(F32(F32(int(src_hw[0])) * s)) + w = int(F32(F32(int(src_hw[1])) * s)) + d = int(pad_divisor) + if d < 1: + raise ValueError("pad_divisor must be >= 1") + return {"src_hw": (int(src_hw[0]), int(src_hw[1])), "resized_hw": (h, w), + "padded_hw": ((h + d - 1) // d * d, (w + d - 1) // d * d)} + + +def _normalize_planes(image: np.ndarray, mean: Any, std: Any) -> np.ndarray: + """BGR uint8 ``(H, W, 3)`` -> float32 planes ``(3, H, W)`` of ``(x - mean_c) / std_c`` as OpenCV evaluates it: + ``convertTo`` with float ``alpha = 1/std``, ``beta = -mean/std`` and one fused multiply-add (exact in float64: + an 8-bit value times a 24-bit alpha plus beta); for alpha == 1 that is the float addition ``x + beta``.""" + img = np.asarray(image) + if img.ndim != 3 or img.shape[2] != 3 or img.dtype != np.uint8: + raise ValueError(f"expected a uint8 (H, W, 3) BGR image, got {img.dtype} {img.shape}") + m = np.asarray(mean, np.float64).reshape(3) + sd = np.asarray(std, np.float64).reshape(3) + alpha = (1.0 / sd).astype(F32) + beta = (-m / sd).astype(F32) + planes = img.transpose(2, 0, 1).astype(F32) # exact: integers 0..255 + for c in range(3): + if alpha[c] == 1.0: + planes[c] += beta[c] # float32, one rounding + else: + planes[c] = planes[c].astype(np.float64) * float(alpha[c]) + float(beta[c]) + return planes + + +def bevformer_normalize(image: np.ndarray, mean: Any = BEVFORMER_MEAN_BGR, std: Any = BEVFORMER_STD) -> np.ndarray: + """One BGR uint8 ``(H, W, 3)`` image -> ``(x - mean_c) / std_c`` as float32 ``(H, W, 3)``, as OpenCV evaluates it + (``convertTo`` with float ``alpha = 1/std``, ``beta = -mean/std``, one fused multiply-add; std 1 deployed).""" + return np.ascontiguousarray(_normalize_planes(image, mean, std).transpose(1, 2, 0)) + + +def bevformer_preprocess(images: Any, *, mean: Any = BEVFORMER_MEAN_BGR, std: Any = BEVFORMER_STD, + scale: float = 0.8, pad_divisor: int = 32, out: Optional[np.ndarray] = None, + workers: int = 1) -> np.ndarray: + """N BGR uint8 camera images (a sequence of ``(H, W, 3)`` or an array ``[N, H, W, 3]``, one source size; the node + feeds 1600x900 after its own ``cv::resize``) -> the network input ``[N, 3, Hp, Wp]`` float32: normalise, resize + x``scale`` (OpenCV float INTER_LINEAR), zero-pad bottom / right to ``pad_divisor``, CHW (B, G, R planes, as + ``to_rgb`` is false). Deployed: ``[6, 3, 736, 1280]`` (the ONNX input ``image`` is this with a leading 1). + ``workers`` > 1 processes the cameras in that many threads (numpy releases the GIL in its loops; about 70 ms per + 1600x900 camera single-threaded on the shared host).""" + imgs = [np.asarray(im) for im in images] + if not imgs: + raise ValueError("no images") + hw = tuple(imgs[0].shape[:2]) + if any(tuple(im.shape[:2]) != hw for im in imgs): + raise ValueError("every camera image must have the same size (the node resizes them to 1600x900 first)") + geo = bevformer_input_geometry(hw, scale, pad_divisor) + (rh, rw), (ph, pw) = geo["resized_hw"], geo["padded_hw"] + lut = opencv_linear_lut(hw, (rh, rw)) # strict: the verified x0.8 geometries + shape = (len(imgs), 3, ph, pw) + if out is None: + out = np.zeros(shape, F32) + else: + if out.shape != shape or out.dtype != F32: + raise ValueError(f"out must be float32 {shape}, got {out.dtype} {out.shape}") + out[:, :, rh:, :] = 0.0 + out[:, :, :, rw:] = 0.0 + def one(i: int) -> None: + lut.apply_planes(_normalize_planes(imgs[i], mean, std), out=out[i, :, :rh, :rw]) + + if int(workers) > 1 and len(imgs) > 1: + from concurrent.futures import ThreadPoolExecutor + + with ThreadPoolExecutor(max_workers=min(int(workers), len(imgs))) as pool: + list(pool.map(one, range(len(imgs)))) + else: + for i in range(len(imgs)): + one(i) + return out + + +# name -> (function, one-line description): the presets this module implements +PRESETS: Dict[str, Tuple[Callable[..., Any], str]] = { + "yolox_letterbox": (yolox_letterbox, "autoware_tensorrt_yolox preprocess.cu:43-129 letterbox (pad 114, BGR)"), + "bevdet_nearest_crop": (bevdet_nearest_crop, "bevdet::Preprocess nearest resize 0.44 + crop (140, 0) of the " + "1600x900 image -> B, G, R uint8 planes"), + "bevformer_preprocess": (bevformer_preprocess, "autoware_tensorrt_bevformer normalise (BGR mean, std 1) -> " + "cv::resize x0.8 INTER_LINEAR on float -> zero pad to /32, CHW"), +} diff --git a/code/tt_diffusion_planner/ttaw/image_area.py b/code/tt_diffusion_planner/ttaw/image_area.py new file mode 100644 index 0000000000000000000000000000000000000000..a3e8385ae18e0a1bf7b9cdea0ddce4b660555a14 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/image_area.py @@ -0,0 +1,179 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C14 image pre-processing, METEOR preset: a bit-exact host emulation of OpenCV's ``cv::resize(..., INTER_AREA)`` on +uint8 images (down-scaling), driven by cached per-size tables. numpy only; no ttnn, no torch, no OpenCV. + +METEOR's runtimes feed every camera as RGB uint8 768x432 and resize larger frames with ``cv2.resize(img, (768, 432), +interpolation=cv2.INTER_AREA)`` (tier4/METEOR ``hf/onnx_smoke_test.py:42-44``; the training extraction +``bevlane/extract_gt.py:114-130`` resized the raw, unrectified JPEGs the same way; research/meteor/SPEC.md section 3), +anisotropically when the aspect ratio differs (2880x1860 -> 768x432). :func:`meteor_resize` reproduces it; the +intrinsics follow with :func:`scale_intrinsics` (row 0 x 768 / w0, row 1 x 432 / h0, ``extract_gt.py:128-130``). + +What OpenCV computes (``modules/imgproc/src/resize.cpp``, identical in 4.8.1 and 4.11.0; the IPP path declines +INTER_AREA because it is not bit-exact, ``ipp_resize``: ``ippSuper`` needs ``useIPP_NotExact``): + +- ``inv_scale = (double)dst / src`` per axis, ``scale = 1.0 / inv_scale`` (double). +- **Integer factors** (``|scale - (int)scale| < DBL_EPSILON`` on both axes; ``resizeAreaFast_``): every output pixel is + the integer sum ``s`` of its ``fx * fy`` source pixels, then ``saturate_cast(s * (1.f / area))``: one float32 + product, rounded half to even (``cvRound``). Exception, the 2x2 ``fast_mode`` of ``ResizeAreaFastVec`` for 1-, 3- + and 4-channel images: ``(s + 2) >> 2`` (half up). Measured on both OpenCV builds (2-channel images take the float + path; 3x3 and 5x5 factors the float path). +- **Other factors** (``resizeArea_``): ``computeResizeAreaTab`` lists, per output column (and row), the + covered source pixels with float32 weights computed in double (``(sx1 - fsx1) / cellWidth``, ``1.0 / cellWidth``, + ``min(min(fsx2 - sx2, 1), cellWidth) / cellWidth``; partial pixels below 1e-3 skipped). Each source row is first + reduced horizontally, ``buf[dx] += S[sx] * alpha`` in table order (float32 multiply, then float32 add: no FMA, the + generic code is compiled for the SSE baseline), then the rows are combined, ``sum[dx] += beta * buf[dx]`` in table + order, and ``saturate_cast(sum)`` rounds half to even and clamps to [0, 255]. + +The emulation replays exactly these float32 operations in the same order, vectorised over the image (zero-weight +padding of the per-pixel tap lists adds exact zeros). Verified against cv2 4.8.1 (tt-env) and 4.11.0 (research venv) +on random and real images of many sizes (``common/tests/host/test_image_area_host.py``): 0 mismatches. +Up-scaling (any axis with ``src < dst``) runs a different OpenCV path (linear-like fixed-point coefficients) and is +refused: METEOR's cameras are all larger than 768x432. +""" +from __future__ import annotations + +import sys +from dataclasses import dataclass +from typing import Any, Optional, Tuple + +import numpy as np + +from .image import LUTCache + +__all__ = ["AreaResizeLUT", "opencv_area_lut", "opencv_resize_area_u8", "meteor_resize", "scale_intrinsics", + "METEOR_INPUT_HW"] + +F32 = np.float32 +METEOR_INPUT_HW = (432, 768) # every METEOR camera slot (meteor_v157.param.yaml:11-12) +_DBL_EPSILON = sys.float_info.epsilon + + +def _area_axis(n_src: int, n_dst: int, scale: float) -> Tuple[np.ndarray, np.ndarray]: + """``computeResizeAreaTab`` of one axis as padded per-output tap lists: (src index [n_dst, T], weight [n_dst, T]) + in table order; unused taps have weight 0 and index 0.""" + taps = [] + for dx in range(n_dst): + fsx1 = dx * scale + fsx2 = fsx1 + scale + cell = min(scale, n_src - fsx1) + sx1 = int(np.ceil(fsx1)) + sx2 = int(np.floor(fsx2)) + sx2 = min(sx2, n_src - 1) + sx1 = min(sx1, sx2) + row = [] + if sx1 - fsx1 > 1e-3: + row.append((sx1 - 1, F32((sx1 - fsx1) / cell))) + for sx in range(sx1, sx2): + row.append((sx, F32(1.0 / cell))) + if fsx2 - sx2 > 1e-3: + row.append((sx2, F32(min(min(fsx2 - sx2, 1.0), cell) / cell))) + taps.append(row) + width = max(len(r) for r in taps) + idx = np.zeros((n_dst, width), np.int64) + w = np.zeros((n_dst, width), F32) + for i, row in enumerate(taps): + for k, (s, a) in enumerate(row): + idx[i, k] = s + w[i, k] = a + return idx, w + + +@dataclass(frozen=True) +class AreaResizeLUT: + """The cached tables of one (source size, destination size) pair; :meth:`apply` resizes one uint8 image.""" + + src_hw: Tuple[int, int] + dst_hw: Tuple[int, int] + fast: Optional[Tuple[int, int]] # integer factors (fy, fx) of the resizeAreaFast_ path, else None + rows: Optional[np.ndarray] # [dst_h, Ty] source rows (general path) + row_w: Optional[np.ndarray] # [dst_h, Ty] float32 beta + cols: Optional[np.ndarray] # [dst_w, Tx] source columns + col_w: Optional[np.ndarray] # [dst_w, Tx] float32 alpha + + def apply(self, image: np.ndarray, *, out: Optional[np.ndarray] = None) -> np.ndarray: + img = np.asarray(image) + if img.dtype != np.uint8 or img.ndim not in (2, 3) or tuple(img.shape[:2]) != self.src_hw: + raise ValueError(f"expected a uint8 {self.src_hw} image (HxW or HxWxC), got {img.dtype} {img.shape}") + squeeze = img.ndim == 2 + x = img[:, :, None] if squeeze else img + dh, dw = self.dst_hw + if self.fast is not None: + fy, fx = self.fast + s = x[:dh * fy, :dw * fx].astype(np.int64).reshape(dh, fy, dw, fx, x.shape[2]).sum(axis=(1, 3)) + if (fy, fx) == (2, 2) and x.shape[2] in (1, 3, 4): + res = (s + 2) >> 2 # ResizeAreaFastVec fast_mode: half up + else: + res = np.rint(s.astype(F32) * F32(1.0 / (fx * fy))) # float32 product, half to even + else: + # horizontal pass per source row: buf[dx] = sum_k S[cols[dx, k]] * col_w[dx, k] (float32, table order) + buf = np.zeros((x.shape[0], dw, x.shape[2]), F32) + src = x.astype(F32) + for k in range(self.cols.shape[1]): + buf += src[:, self.cols[:, k], :] * self.col_w[None, :, k, None] + # vertical pass: sum[dy] = sum_j beta[dy, j] * buf[rows[dy, j]] (float32, table order) + acc = np.zeros((dh, dw, x.shape[2]), F32) + for j in range(self.rows.shape[1]): + acc += self.row_w[:, j, None, None] * buf[self.rows[:, j]] + res = np.rint(acc) # cvRound: half to even + res = np.clip(res, 0, 255).astype(np.uint8) + if squeeze: + res = res[:, :, 0] + if out is not None: + out[...] = res + return out + return res + + +_AREA_LUTS = LUTCache(maxsize=16) # 8 cameras of up to two sizes each + + +def opencv_area_lut(src_hw: Tuple[int, int], dst_hw: Tuple[int, int]) -> AreaResizeLUT: + """Tables of ``cv::resize(INTER_AREA)`` from ``src_hw`` down to ``dst_hw`` (cached per size pair).""" + src_hw = (int(src_hw[0]), int(src_hw[1])) + dst_hw = (int(dst_hw[0]), int(dst_hw[1])) + return _AREA_LUTS.get((src_hw, dst_hw), lambda: _build_area_lut(src_hw, dst_hw)) + + +def _build_area_lut(src_hw: Tuple[int, int], dst_hw: Tuple[int, int]) -> AreaResizeLUT: + (sh, sw), (dh, dw) = src_hw, dst_hw + if min(sh, sw, dh, dw) < 1: + raise ValueError(f"bad sizes {src_hw} -> {dst_hw}") + if dh > sh or dw > sw: + raise ValueError(f"INTER_AREA up-scaling {src_hw} -> {dst_hw} is not emulated (OpenCV uses another path)") + scale_y = 1.0 / (float(dh) / sh) + scale_x = 1.0 / (float(dw) / sw) + iy, ix = int(round(scale_y)), int(round(scale_x)) + if abs(scale_x - ix) < _DBL_EPSILON and abs(scale_y - iy) < _DBL_EPSILON: + return AreaResizeLUT(src_hw, dst_hw, (iy, ix), None, None, None, None) + rows, row_w = _area_axis(sh, dh, scale_y) + cols, col_w = _area_axis(sw, dw, scale_x) + return AreaResizeLUT(src_hw, dst_hw, None, rows, row_w, cols, col_w) + + +def opencv_resize_area_u8(image: np.ndarray, dst_hw: Tuple[int, int], *, + out: Optional[np.ndarray] = None) -> np.ndarray: + """``cv2.resize(image, (dst_w, dst_h), interpolation=cv2.INTER_AREA)`` for uint8 HxW / HxWxC down-scaling; + the same size returns a copy (OpenCV's ``copyTo`` shortcut).""" + img = np.asarray(image) + if tuple(img.shape[:2]) == tuple(dst_hw): + res = np.array(img, dtype=np.uint8, copy=True) + if out is not None: + out[...] = res + return out + return res + return opencv_area_lut(img.shape[:2], dst_hw).apply(img, out=out) + + +def meteor_resize(image: np.ndarray, *, out: Optional[np.ndarray] = None) -> np.ndarray: + """One camera frame (uint8 HxWx3, any channel order: the resize is per channel) -> 432x768 as METEOR's runtimes + (``INTER_AREA`` when the size differs, a copy otherwise).""" + return opencv_resize_area_u8(image, METEOR_INPUT_HW, out=out) + + +def scale_intrinsics(K: Any, src_hw: Tuple[int, int], dst_hw: Tuple[int, int] = METEOR_INPUT_HW) -> np.ndarray: + """K of a ``src_hw`` image -> K of the resized ``dst_hw`` image: row 0 x dst_w / src_w, row 1 x dst_h / src_h + (``extract_gt.py:128-130``; skew is scaled with row 0 and ignored by the network). float64.""" + k = np.array(K, dtype=np.float64).reshape(3, 3).copy() + k[0] *= dst_hw[1] / float(src_hw[1]) + k[1] *= dst_hw[0] / float(src_hw[0]) + return k diff --git a/code/tt_diffusion_planner/ttaw/image_linear.py b/code/tt_diffusion_planner/ttaw/image_linear.py new file mode 100644 index 0000000000000000000000000000000000000000..d31a987cd97b797f19977aba4e3c4e95a0e6143e --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/image_linear.py @@ -0,0 +1,292 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C14 image pre-processing, SceneSeg preset: bit-exact host emulations of OpenCV's ``cv::resize`` on uint8 images +with ``INTER_LINEAR`` (the default interpolation) and ``INTER_NEAREST``, driven by cached per-size tables, and the +pre-processing of Autoware VisionPilot's SceneSeg node built on them. numpy only; no ttnn, no torch, no OpenCV. + +The SceneSeg node (VisionPilot ``middleware_recipes/common/backends/onnx_runtime_backend.cpp:41-60`` and the identical +``tensorrt_backend.cpp:160-177`` at autowarefoundation/vision_pilot ``04fa3e80``; research/sceneseg/SPEC.md section 3) +takes the ``bgr8`` image of ``cv_bridge``, then:: + + cv::resize(input_image, resized_image, cv::Size(640, 320)); // INTER_LINEAR, aspect ratio not kept + resized_image.convertTo(float_image, CV_32FC3, 1.0 / 255.0); + cv::subtract(float_image, cv::Scalar(0.406, 0.456, 0.485), float_image); + cv::divide(float_image, cv::Scalar(0.225, 0.224, 0.229), float_image); + cv::split(float_image, channels); // planes B, G, R -> NCHW + +and publishes the argmax mask resized back to the source size with ``cv::resize(..., INTER_NEAREST)`` +(``run_model_node.cpp:117-188``). :func:`sceneseg_resize` / :func:`sceneseg_normalize` / :func:`sceneseg_preprocess` +reproduce the first part, :func:`opencv_resize_nearest` the mask resize. + +What OpenCV computes (``modules/imgproc/src/resize.cpp`` and ``core/src/arithm.cpp``; identical in 4.8.1 and 4.11.0 +on this x86-64 host, measured; the IPP path declines both resizes: uint8 ``ippLinear`` and ``ippNearest`` need +``useIPP_NotExact``): + +- ``cv::resize`` with ``dsize == ssize`` copies. ``inv_scale = (double)dst / src`` and ``scale = 1.0 / inv_scale`` + per axis. +- **INTER_LINEAR, exactly 2x down on both axes**: OpenCV switches to ``INTER_AREA`` (``resizeAreaFast_``), which + :func:`.image_area.opencv_resize_area_u8` emulates (2x2 ``fast_mode``: ``(s + 2) >> 2``). +- **INTER_LINEAR otherwise** (``resizeGeneric_`` with ``HResizeLinear`` and the uchar + specialisation of ``VResizeLinear``): per output column ``fx = (float)((dx + 0.5) * scale - 0.5)``, + ``sx = floor(fx)``, ``fx -= sx`` (float); at the borders ``sx < 0 -> sx = 0, fx = 0`` and ``sx >= W - 1 -> + sx = W - 1, fx = 0``; integer weights ``a0 = cvRound((1.f - fx) * 2048)``, ``a1 = cvRound(fx * 2048)`` (round half + to even). Rows take the same coefficients **without** the border fix (the row indices are clipped to [0, H - 1] + instead, ``clip(sy + k, 0, H)``). The horizontal pass is exact integer arithmetic, + ``h = S[sx] * a0 + S[sx + 1] * a1`` (int32), and the vertical pass is the fixed-point + ``(((b0 * (h0 >> 4)) >> 16) + ((b1 * (h1 >> 4)) >> 16) + 2) >> 2`` that the SIMD and the scalar code share. +- **INTER_NEAREST** (``resizeNN``): ``sx = min(floor(dx * (1.0 / inv_scale_x)), W - 1)`` in double, rows alike. +- **Normalisation**: ``convertTo(CV_32F, 1/255)`` is one float32 product ``x * (float)(1/255)``; ``cv::subtract`` + with a ``Scalar`` runs in float32 (``x - (float)mean_c``); ``cv::divide`` with a ``Scalar`` runs in double and + rounds once to float32 (``(float)((double)x / std_c)``), measured with OpenCV's own scalar arithmetic. A float32 + division instead differs by one ulp in about 27 % of the values (the research reference's arithmetic). + +Channel orders (SPEC section 3, PLAN.md D10): ``"bgr"`` is the deployed path: B, G, R planes normalised with the +BGR-ordered ImageNet constants, fed to a network trained on R, G, B. ``"rgb"`` is the training-consistent option: the +same arithmetic on R, G, B planes with the RGB constants (an exact channel permutation of the image first, as a +``cv_bridge`` ``rgb8`` conversion would give). Every physical colour gets the same mean and std in both orders, so +the two inputs are channel permutations of each other. + +Verified bit-exact against cv2 4.8.1 (tt-env) and 4.11.0 (research venv) on random and real images of many sizes +(down- and up-scaling, odd sizes, 1-4 channels), against a scalar transliteration of ``resize.cpp``, and against the +research golden inputs of the SceneSeg port (``common/tests/host/test_image_sceneseg_host.py``). +""" +from __future__ import annotations + +from dataclasses import dataclass +from typing import Any, Optional, Tuple + +import numpy as np + +from .image import LUTCache +from .image_area import opencv_resize_area_u8 + +__all__ = [ + "LinearResizeU8LUT", + "opencv_linear_u8_lut", + "opencv_resize_linear_u8", + "NearestResizeLUT", + "opencv_nearest_lut", + "opencv_resize_nearest", + "sceneseg_resize", + "sceneseg_normalize", + "sceneseg_preprocess", + "SCENESEG_INPUT_HW", + "SCENESEG_MEAN_RGB", + "SCENESEG_STD_RGB", + "SCENESEG_CHANNEL_ORDERS", +] + +F32 = np.float32 +INTER_RESIZE_COEF_SCALE = 2048 # 1 << INTER_RESIZE_COEF_BITS (11), resize.cpp +SCENESEG_INPUT_HW = (320, 640) # (H, W) of the ONNX input; fixed by the SceneContext reshape (SPEC 4.2) +SCENESEG_MEAN_RGB = (0.485, 0.456, 0.406) # ImageNet, R G B (the C++ code lists them reversed for B G R planes) +SCENESEG_STD_RGB = (0.229, 0.224, 0.225) +SCENESEG_CHANNEL_ORDERS = ("bgr", "rgb") + + +def _check_hw(hw: Tuple[int, int], what: str) -> Tuple[int, int]: + h, w = (int(v) for v in hw) + if h < 1 or w < 1: + raise ValueError(f"{what} must be positive, got {hw}") + return h, w + + +def _as_image(image: Any, dtype: Optional[type] = np.uint8) -> np.ndarray: + a = np.asarray(image) + if a.ndim not in (2, 3) or a.shape[0] < 1 or a.shape[1] < 1: + raise ValueError(f"expected an HxW or HxWxC image, got shape {a.shape}") + if dtype is not None and a.dtype != dtype: + raise ValueError(f"expected {np.dtype(dtype).name} data, got {a.dtype}") + return a + + +# ----------------------------------------------------------------------------------------------- INTER_LINEAR + +def _linear_axis(n_src: int, n_dst: int, border_fix: bool) -> Tuple[np.ndarray, np.ndarray, np.ndarray, np.ndarray]: + """One axis of ``resize()``'s INTER_LINEAR tables for uint8: source indices (clipped) and integer weights.""" + scale = 1.0 / (n_dst / n_src) # scale = 1. / inv_scale, inv = (double)dst / src + f = ((np.arange(n_dst, dtype=np.float64) + 0.5) * scale - 0.5).astype(F32) + s = np.floor(f).astype(np.int64) # cvFloor(fx) + f = (f - s.astype(F32)).astype(F32) # fx -= sx (float; exact) + if border_fix: # x axis only (the y loop has no such fix) + lo = s < 0 + f[lo], s[lo] = F32(0), 0 + hi = s >= n_src - 1 + f[hi], s[hi] = F32(0), n_src - 1 + one = F32(INTER_RESIZE_COEF_SCALE) + w0 = np.rint((F32(1) - f).astype(F32) * one).astype(np.int32) # saturate_cast(cbuf[k] * 2048): cvRound + w1 = np.rint(f * one).astype(np.int32) + return np.clip(s, 0, n_src - 1), np.clip(s + 1, 0, n_src - 1), w0, w1 + + +@dataclass(frozen=True) +class LinearResizeU8LUT: + """The tables of one ``cv::resize(uint8, INTER_LINEAR)`` geometry. ``mode``: ``"copy"`` (same size), + ``"area2x"`` (exact 2x down on both axes: OpenCV's INTER_AREA fast path) or ``"linear"``.""" + + src_hw: Tuple[int, int] + dst_hw: Tuple[int, int] + mode: str + cols0: Optional[np.ndarray] = None # (w,) source column of the first tap (clipped) + cols1: Optional[np.ndarray] = None # (w,) second tap + a0: Optional[np.ndarray] = None # (w,) int32 weights, a0 + a1 == 2048 up to OpenCV's rounding + a1: Optional[np.ndarray] = None + rows: Optional[np.ndarray] = None # (R,) the source rows the vertical pass reads (sorted, unique) + r0: Optional[np.ndarray] = None # (h,) index into ``rows`` of the first row tap + r1: Optional[np.ndarray] = None # (h,) second row tap + b0: Optional[np.ndarray] = None # (h,) int32 row weights + b1: Optional[np.ndarray] = None + + def apply(self, image: Any, *, out: Optional[np.ndarray] = None) -> np.ndarray: + """Resize one HxW or HxWxC uint8 image of size ``src_hw`` -> ``dst_hw`` (same channel count).""" + img = _as_image(image) + if img.shape[:2] != self.src_hw: + raise ValueError(f"table built for {self.src_hw}, image is {img.shape[:2]}") + if self.mode == "copy": + res = img.copy() + elif self.mode == "area2x": + res = opencv_resize_area_u8(img, self.dst_hw) + else: + src = img if img.ndim == 3 else img[:, :, None] + sub = src[self.rows] # only the rows the vertical pass reads + h = sub[:, self.cols0, :].astype(np.int32) # horizontal pass, exact int32 (<= 255 * 2048) + h *= self.a0[None, :, None] + h1 = sub[:, self.cols1, :].astype(np.int32) + h1 *= self.a1[None, :, None] + h += h1 + h >>= 4 # S[x] >> 4 (h >= 0) + v = h[self.r0] # vertical pass, the uchar VResizeLinear + v *= self.b0[:, None, None] + v >>= 16 + v1 = h[self.r1] + v1 *= self.b1[:, None, None] + v1 >>= 16 + v += v1 + v += 2 + v >>= 2 + res = v.astype(np.uint8) # in [0, 255] (weights sum to 2048) + if img.ndim == 2: + res = res[:, :, 0] + if out is not None: + out[...] = res + return out + return res + + +def _build_linear_lut(src_hw: Tuple[int, int], dst_hw: Tuple[int, int]) -> LinearResizeU8LUT: + (H, W), (h, w) = src_hw, dst_hw + if (h, w) == (H, W): + return LinearResizeU8LUT(src_hw, dst_hw, "copy") + if 2 * h == H and 2 * w == W: # is_area_fast && iscale_x == iscale_y == 2 + return LinearResizeU8LUT(src_hw, dst_hw, "area2x") + c0, c1, a0, a1 = _linear_axis(W, w, True) + y0, y1, b0, b1 = _linear_axis(H, h, False) + rows = np.unique(np.concatenate([y0, y1])) + r0, r1 = np.searchsorted(rows, y0), np.searchsorted(rows, y1) + return LinearResizeU8LUT(src_hw, dst_hw, "linear", c0, c1, a0, a1, rows, r0, r1, b0, b1) + + +_LINEAR_U8_LUTS = LUTCache(maxsize=8) + + +def opencv_linear_u8_lut(src_hw: Tuple[int, int], dst_hw: Tuple[int, int]) -> LinearResizeU8LUT: + """The cached tables of ``cv::resize(uint8 src_hw -> dst_hw, INTER_LINEAR)`` (any scale, up or down).""" + key = (_check_hw(src_hw, "src_hw"), _check_hw(dst_hw, "dst_hw")) + return _LINEAR_U8_LUTS.get(key, lambda: _build_linear_lut(*key)) + + +def opencv_resize_linear_u8(image: Any, dst_hw: Tuple[int, int], *, out: Optional[np.ndarray] = None) -> np.ndarray: + """``cv::resize(image, dst, Size(w, h))`` (INTER_LINEAR) of an HxW or HxWxC uint8 image, bit-exact.""" + img = _as_image(image) + return opencv_linear_u8_lut(img.shape[:2], dst_hw).apply(img, out=out) + + +# ---------------------------------------------------------------------------------------------- INTER_NEAREST + +@dataclass(frozen=True) +class NearestResizeLUT: + """``resizeNN`` index tables: ``rows`` (h,), ``cols`` (w,); a copy when the sizes are equal.""" + + src_hw: Tuple[int, int] + dst_hw: Tuple[int, int] + rows: np.ndarray + cols: np.ndarray + + def apply(self, image: Any, *, out: Optional[np.ndarray] = None) -> np.ndarray: + """Resize one HxW or HxWxC image of any dtype (a gather: values are copied, never mixed).""" + img = _as_image(image, dtype=None) + if img.shape[:2] != self.src_hw: + raise ValueError(f"table built for {self.src_hw}, image is {img.shape[:2]}") + res = img.copy() if self.src_hw == self.dst_hw else img[self.rows][:, self.cols] + if out is not None: + out[...] = res + return out + return res + + +def _nearest_axis(n_src: int, n_dst: int) -> np.ndarray: + inv = 1.0 / (n_dst / n_src) # ifx = 1. / fx, fx = inv_scale_x + return np.minimum(np.floor(np.arange(n_dst, dtype=np.float64) * inv).astype(np.int64), n_src - 1) + + +_NEAREST_LUTS = LUTCache(maxsize=8) + + +def opencv_nearest_lut(src_hw: Tuple[int, int], dst_hw: Tuple[int, int]) -> NearestResizeLUT: + """The cached index tables of ``cv::resize(src_hw -> dst_hw, INTER_NEAREST)``.""" + key = (_check_hw(src_hw, "src_hw"), _check_hw(dst_hw, "dst_hw")) + return _NEAREST_LUTS.get(key, lambda: NearestResizeLUT(key[0], key[1], _nearest_axis(key[0][0], key[1][0]), + _nearest_axis(key[0][1], key[1][1]))) + + +def opencv_resize_nearest(image: Any, dst_hw: Tuple[int, int], *, out: Optional[np.ndarray] = None) -> np.ndarray: + """``cv::resize(image, dst, Size(w, h), 0, 0, INTER_NEAREST)`` of an HxW or HxWxC image, bit-exact.""" + img = _as_image(image, dtype=None) + return opencv_nearest_lut(img.shape[:2], dst_hw).apply(img, out=out) + + +# ------------------------------------------------------------------------------------------- SceneSeg preset + +def _order(channel_order: str) -> str: + if channel_order not in SCENESEG_CHANNEL_ORDERS: + raise ValueError(f"channel_order must be one of {SCENESEG_CHANNEL_ORDERS}, got {channel_order!r}") + return channel_order + + +def sceneseg_resize(image_bgr: Any, *, dst_hw: Tuple[int, int] = SCENESEG_INPUT_HW, + out: Optional[np.ndarray] = None) -> np.ndarray: + """The node's squash, ``cv::resize(input_image, resized, Size(640, 320))`` (INTER_LINEAR, aspect ratio not + kept) of one HxWx3 uint8 image (channel order kept: give it B, G, R as ``cv_bridge`` ``bgr8`` does) -> + ``(320, 640, 3)`` uint8. This is the device input of the TT port (614 KB).""" + img = _as_image(image_bgr) + if img.ndim != 3 or img.shape[2] != 3: + raise ValueError(f"expected an HxWx3 BGR image, got shape {img.shape}") + return opencv_resize_linear_u8(img, dst_hw, out=out) + + +def sceneseg_normalize(resized_bgr: Any, *, channel_order: str = "bgr") -> np.ndarray: + """The node's normalisation of the resized B, G, R uint8 image -> the ONNX input ``float32 [1, 3, H, W]``. + + ``"bgr"`` (deployed): planes B, G, R; ``convertTo(CV_32F, 1/255)`` (float32 product), ``cv::subtract`` of + (0.406, 0.456, 0.485) in float32, ``cv::divide`` by (0.225, 0.224, 0.229) in double rounded to float32. + ``"rgb"`` (training order): the same arithmetic on planes R, G, B with (0.485, 0.456, 0.406) / + (0.229, 0.224, 0.225).""" + img = _as_image(resized_bgr) + if img.ndim != 3 or img.shape[2] != 3: + raise ValueError(f"expected an HxWx3 BGR image, got shape {img.shape}") + if _order(channel_order) == "rgb": + img = img[:, :, ::-1] + mean, std = SCENESEG_MEAN_RGB, SCENESEG_STD_RGB + else: + mean, std = SCENESEG_MEAN_RGB[::-1], SCENESEG_STD_RGB[::-1] + planes = np.ascontiguousarray(img.transpose(2, 0, 1)).astype(F32) # cv::split -> NCHW + x = planes * F32(1.0 / 255.0) # convertTo: float32 product + x = x - np.asarray(mean, F32)[:, None, None] # cv::subtract (float32) + x = (x.astype(np.float64) / np.asarray(std, np.float64)[:, None, None]).astype(F32) # cv::divide (double) + return x[None] + + +def sceneseg_preprocess(image_bgr: Any, *, channel_order: str = "bgr") -> np.ndarray: + """One HxWx3 B, G, R uint8 camera image -> the SceneSeg ONNX input ``float32 [1, 3, 320, 640]``, bit-exact with + the deployed C++ (``"bgr"``) or its training-order variant (``"rgb"``): :func:`sceneseg_normalize` of + :func:`sceneseg_resize`.""" + return sceneseg_normalize(sceneseg_resize(image_bgr), channel_order=channel_order) diff --git a/code/tt_diffusion_planner/ttaw/image_triangle.py b/code/tt_diffusion_planner/ttaw/image_triangle.py new file mode 100644 index 0000000000000000000000000000000000000000..842525442ba06b48d6909df3bf68f9026c6b9d89 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/image_triangle.py @@ -0,0 +1,344 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C14 image pre-processing, StreamPETR preset: a bit-exact host emulation of Autoware's fused anti-aliased resize + +crop + normalise CUDA kernel (a PIL-style triangle filter with adaptive support), driven by cached per-size tables. +numpy only; no ttnn, no torch, no OpenCV. + +Deployed code (autoware_universe main @ 9ceaccf, ``perception/autoware_camera_streampetr``; research/streampetr/SPEC.md +section 3.1): + +- ``calculate_image_processing_params`` (``lib/network/camera_data_store.cpp:287-316``), float32: + ``resize = max(480 / (float)H, 640 / (float)W)``; ``new_w = (int)(W * resize)``, ``new_h = (int)(H * resize)``; + ``crop_h = max(0, (int)(1.0f * new_h) - 480)`` (the top rows are cropped, the bottom kept: ``bottom_crop_portion`` + 0), ``crop_w = max(0, (new_w - 640) / 2)`` (centred) -> :func:`autoware_resize_geometry`. +- ``resize_and_extract_roi_kernel`` (``lib/network/preprocess_kernel.cu:73-176``; the 0.52.1 kernel + ``resizeAndExtractRoi_kernel`` has the same arithmetic), per output pixel, all float32: + ``r = (float)(out + roi_start)``, ``scale = (float)in / (float)resized``, ``centre = (r + 0.5f) * scale - 0.5f``, + ``support = scale > 1 ? scale : 1``, taps ``ceilf(centre - support) .. floorf(centre + support)`` clamped to the + image, ``w = max(0, 1 - |centre - i| / support)`` per axis, ``w = wx * wy``; the loop runs **y outer, x inner** + and skips ``w <= 0``: ``sum_c += pix_c * w``, ``sum_w += w``; then ``sum_c /= sum_w`` (when ``sum_w > 0``) and + ``out_c = (sum_c - mean[c]) / std[c]``, written planar (CHW). The channel order of the output is the byte order + of the buffer the kernel reads: RGB on main (a ``bgr8`` message is swapped in place first, + ``convert_bgr_to_rgb_kernel``), the message order on 0.52.1. + +The emulation replays exactly these float32 operations in the same order, vectorised over the output pixels: per +output row / column the source taps and the 1-D weights are tabulated once (:class:`TriangleResizeLUT`), and the 2-D +accumulation runs tap pair by tap pair in the kernel's (y outer, x inner) order. Taps outside a pixel's window carry +weight 0, which adds exact zeros (the kernel skips them). + +**Floating-point contraction** (``fma``). Written as C, every operation rounds once: ``fma=False`` (default) does +that, and equals the SPEC's numpy port (``research/streampetr/scripts/sp_common.py``) and the research goldens bit +for bit (their ``inputs_sha256.json``). The node's CMakeLists passes no ``--fmad`` / ``--use_fast_math`` flag, so +nvcc compiles with its default ``--fmad=true``, which lets the compiler fuse a multiply feeding an add into one +``fmaf``. ``fma=True`` models LLVM NVPTX's aggressive fusion: ``centre = fmaf(r + 0.5f, scale, -0.5f)``, +``sum_c = fmaf(pix_c, w, sum_c)`` and ``sum_w = fmaf(wx, wy, sum_w)`` (the rounded ``w`` still feeds the channel +products). Which form the deployed binary has cannot be checked without CUDA. The two differ in ~43 % of the +inputs by at most ~3e-5 in normalised units (one ulp of a ~1000 px centre moves every tap weight; measured in +``common/tests/host/test_image_triangle_host.py``), two orders of magnitude below the bf16 rounding of the device +input. + +Normalisation presets (:data:`STREAMPETR_PRESETS`, SPEC section 3.4; PLAN.md D7 picks ``autoware_main``): + +================== ====================== ================================================================== +preset planes (buffer order) statistics +================== ====================== ================================================================== +``autoware_main`` R, G, B mean (123.675, 116.28, 103.53), std (58.395, 57.12, 57.375) + (universe main / 0.53.0, ``camera_data_store.cpp:128-134``) +``autoware_0.52`` R, G, B (an ``rgb8`` mean (103.53, 116.28, 123.675), std (57.375, 57.12, 58.395) on the + camera, as TIER IV's) message bytes (0.52.1, ``camera_data_store.cpp:122-127``) +``awml_training`` B, G, R the 0.52 statistics (AWML t4base v2.5 test pipeline: cv2 decode, + ``to_rgb=False``; = 0.52.1 fed a ``bgr8`` camera) +================== ====================== ================================================================== +""" +from __future__ import annotations + +from concurrent.futures import ThreadPoolExecutor +from dataclasses import dataclass +from typing import Any, Dict, Optional, Sequence, Tuple + +import numpy as np + +from .image import LUTCache + +__all__ = [ + "TriangleGeometry", + "TriangleResizeLUT", + "StreamPETRPreset", + "autoware_resize_geometry", + "triangle_resize_lut", + "fma_f32", + "streampetr_preprocess", + "streampetr_preprocess_batch", + "STREAMPETR_PRESETS", + "STREAMPETR_INPUT_HW", + "STREAMPETR_MEAN_RGB", + "STREAMPETR_STD_RGB", + "STREAMPETR_MEAN_BGR", + "STREAMPETR_STD_BGR", +] + +F32 = np.float32 +F64 = np.float64 +STREAMPETR_INPUT_HW = (480, 640) # HF ml_package_camera_streampetr.param.yaml:7-8 (input_image_height / _width) +# camera_data_store.cpp:128-134 (main: RGB order) and AW052 camera_data_store.cpp:122-127 (0.52.1: BGR order) +STREAMPETR_MEAN_RGB = (123.675, 116.280, 103.530) +STREAMPETR_STD_RGB = (58.395, 57.120, 57.375) +STREAMPETR_MEAN_BGR = (103.530, 116.280, 123.675) +STREAMPETR_STD_BGR = (57.375, 57.120, 58.395) + + +@dataclass(frozen=True) +class StreamPETRPreset: + """One normalisation preset: the plane order the kernel sees (``"rgb"`` or ``"bgr"``) and the per-plane + statistics, as float32 values (the kernel reads them from a float device buffer).""" + + name: str + planes: str + mean: Tuple[float, float, float] + std: Tuple[float, float, float] + description: str = "" + + def mean_f32(self) -> np.ndarray: + return np.asarray(self.mean, dtype=F32) + + def std_f32(self) -> np.ndarray: + return np.asarray(self.std, dtype=F32) + + +STREAMPETR_PRESETS: Dict[str, StreamPETRPreset] = { + "autoware_main": StreamPETRPreset("autoware_main", "rgb", STREAMPETR_MEAN_RGB, STREAMPETR_STD_RGB, + "Autoware universe main / 0.53.0: RGB planes, RGB-ordered statistics"), + "autoware_0.52": StreamPETRPreset("autoware_0.52", "rgb", STREAMPETR_MEAN_BGR, STREAMPETR_STD_BGR, + "Autoware 0.52.1 fed an rgb8 camera: message-order (RGB) planes, " + "BGR-ordered statistics"), + "awml_training": StreamPETRPreset("awml_training", "bgr", STREAMPETR_MEAN_BGR, STREAMPETR_STD_BGR, + "AWML t4base v2.5 training / test pipeline (= 0.52.1 fed a bgr8 camera): " + "BGR planes, BGR statistics"), +} + + +# --------------------------------------------------------------------------------------------- geometry + + +@dataclass(frozen=True) +class TriangleGeometry: + """``calculate_image_processing_params`` of one source size: the virtual resized image and the ROI cut from it.""" + + src_hw: Tuple[int, int] + resized_hw: Tuple[int, int] # (new_h, new_w) + roi_hw: Tuple[int, int] # the network input (480, 640) + roi_start: Tuple[int, int] # (start_y, start_x) in the resized image + resize: float # the float32 resize factor, as an exact Python float + + @property + def resize_f32(self) -> np.float32: + return F32(self.resize) + + def to_dict(self) -> Dict[str, Any]: + return {"src_hw": list(self.src_hw), "resized_hw": list(self.resized_hw), "roi_hw": list(self.roi_hw), + "roi_start": list(self.roi_start), "resize": self.resize} + + +def autoware_resize_geometry(src_h: int, src_w: int, dst_h: int = STREAMPETR_INPUT_HW[0], + dst_w: int = STREAMPETR_INPUT_HW[1]) -> TriangleGeometry: + """``camera_data_store.cpp:287-316`` in float32 (module docstring).""" + src_h, src_w, dst_h, dst_w = int(src_h), int(src_w), int(dst_h), int(dst_w) + if min(src_h, src_w, dst_h, dst_w) <= 0: + raise ValueError(f"sizes must be positive: src {src_h}x{src_w}, dst {dst_h}x{dst_w}") + scale_h = F32(F32(dst_h) / F32(src_h)) + scale_w = F32(F32(dst_w) / F32(src_w)) + resize = scale_w if scale_h < scale_w else scale_h # std::max(scaleH, scaleW) + new_w = int(F32(F32(src_w) * resize)) # static_cast(int * float): truncation + new_h = int(F32(F32(src_h) * resize)) + crop_h = max(0, int(F32(F32(1.0) * F32(new_h))) - dst_h) # (1.0f - bottom_crop_portion) * newH, portion 0 + crop_w = max(0, int((new_w - dst_w) / 2)) if new_w > dst_w else 0 # C int division, then std::max(0, .) + return TriangleGeometry((src_h, src_w), (new_h, new_w), (dst_h, dst_w), (max(0, crop_h), max(0, crop_w)), + float(resize)) + + +# --------------------------------------------------------------------------------------------- arithmetic + + +def fma_f32(a: Any, b: Any, c: Any) -> np.ndarray: + """``fmaf(a, b, c)`` for float32 arrays: ``a * b + c`` rounded ONCE to float32 (round to nearest even), exactly. + + The float64 product of two float32 values is exact; the float64 sum may round, and a float64 result that lands + on a float32 tie the exact value does not sit on would round twice. The TwoSum error term detects that case and + the tie is broken towards the exact value.""" + a64 = np.asarray(a, dtype=F32).astype(F64) + b64 = np.asarray(b, dtype=F32).astype(F64) + c64 = np.asarray(c, dtype=F32).astype(F64) + p = a64 * b64 + s = p + c64 + r = np.asarray(s.astype(F32)) + bv = s - p + err = (p - (s - bv)) + (c64 - bv) # p + c == s + err exactly (Knuth's TwoSum) + bad = np.nonzero(np.broadcast_to(err, s.shape) != 0) + if bad[0].size: + sb, eb = s[bad], np.broadcast_to(err, s.shape)[bad] + rb = r[bad] + d = sb - rb.astype(F64) + nb = np.nextafter(rb, np.where(d > 0, np.inf, -np.inf).astype(F32)) + tie = (d != 0) & (sb == (rb.astype(F64) + nb.astype(F64)) * 0.5) + fix = tie & (np.sign(eb) == np.sign(d)) + rb = np.where(fix, nb, rb) + r = r.copy() + r[bad] = rb + return r + + +def _axis_taps(n_out: int, start: int, n_in: int, n_resized: int, fma: bool) -> Tuple[np.ndarray, np.ndarray]: + """Source index [n_out, T] and float32 weight [n_out, T] of every output row (or column), in the kernel's tap + order; taps past a pixel's window have weight 0 (index clamped into the image).""" + scale = F32(F32(n_in) / F32(n_resized)) + support = scale if scale > F32(1.0) else F32(1.0) + r = (np.arange(n_out, dtype=np.int64) + int(start)).astype(F32) # (float)(out + roi_start), exact + r5 = (r + F32(0.5)).astype(F32) + if fma: + centre = (r5.astype(F64) * F64(scale) - 0.5).astype(F32) # one rounding (exact in float64) + else: + centre = ((r5 * scale).astype(F32) - F32(0.5)).astype(F32) + lo = np.maximum(0, np.ceil((centre - support).astype(F32)).astype(np.int64)) + hi = np.minimum(n_in - 1, np.floor((centre + support).astype(F32)).astype(np.int64)) + n_taps = max(1, int((hi - lo).max()) + 1) + idx = lo[:, None] + np.arange(n_taps, dtype=np.int64)[None, :] + valid = idx <= hi[:, None] + idx = np.minimum(idx, n_in - 1) + d = (centre[:, None] - idx.astype(F32)).astype(F32) + w = (F32(1.0) - (np.abs(d) / support).astype(F32)).astype(F32) + w = np.maximum(F32(0.0), w).astype(F32) + w = np.where(valid, w, F32(0.0)).astype(F32) + return idx, w + + +@dataclass(frozen=True) +class TriangleResizeLUT: + """The per-size tables of ``resize_and_extract_roi_kernel``: build with :func:`triangle_resize_lut` (cached), + apply with :meth:`apply` (the per-pixel weighted mean, before the mean / std normalisation).""" + + geometry: TriangleGeometry + rows: np.ndarray # (roi_h, Ty) int64 source rows + row_w: np.ndarray # (roi_h, Ty) float32 wy (0 past the window) + cols: np.ndarray # (roi_w, Tx) int64 source columns + col_w: np.ndarray # (roi_w, Tx) float32 wx + fma: bool + + def apply(self, image: np.ndarray, *, out: Optional[np.ndarray] = None) -> np.ndarray: + """One uint8 image ``(src_h, src_w, C)`` (any channel count, channel order kept) -> float32 ``(C, roi_h, + roi_w)``: ``sum_c / sum_w`` (``sum_c`` where ``sum_w == 0``), exactly as the kernel before step 9.""" + g = self.geometry + img = np.asarray(image) + if img.ndim != 3 or img.dtype != np.uint8 or tuple(img.shape[:2]) != g.src_hw: + raise ValueError(f"expected a uint8 ({g.src_hw[0]}, {g.src_hw[1]}, C) image, got {img.dtype} {img.shape}") + roi_h, roi_w = g.roi_hw + cn = img.shape[2] + acc = np.zeros((cn, roi_h, roi_w), F32) + wsum = np.zeros((roi_h, roi_w), F32) + pix = np.empty((cn, roi_h, roi_w), F32) + planes = np.ascontiguousarray(img.transpose(2, 0, 1)) # (C, H, W) uint8 + # the source columns of every horizontal tap, once per image (float32 holds uint8 exactly) + cols = [planes[:, :, self.cols[:, tx]].astype(F32) for tx in range(self.cols.shape[1])] + for ty in range(self.rows.shape[1]): # y outer ... + rows = self.rows[:, ty] + wy = self.row_w[:, ty][:, None] + for tx in range(self.cols.shape[1]): # ... x inner, as the kernel + wx = self.col_w[:, tx][None, :] + w = (wx * wy).astype(F32) # w = wx * wy (rounded) + np.take(cols[tx], rows, axis=1, out=pix) # (C, roi_h, roi_w) pixels + if self.fma: + acc = fma_f32(pix, w, acc) # sum_c = fmaf(pix, w, sum_c) + wsum = fma_f32(np.broadcast_to(wx, w.shape), np.broadcast_to(wy, w.shape), wsum) + else: + np.multiply(pix, w, out=pix) # pix * w (rounded) + np.add(acc, pix, out=acc) # sum_c + . (rounded) + np.add(wsum, w, out=wsum) + pos = wsum > 0 + result = np.where(pos, acc / np.where(pos, wsum, F32(1.0)), acc).astype(F32) + if out is not None: + if out.shape != result.shape or out.dtype != F32: + raise ValueError(f"out must be float32 {result.shape}, got {out.dtype} {out.shape}") + out[...] = result + return out + return result + + +_TRIANGLE_LUTS = LUTCache(maxsize=16) + + +def triangle_resize_lut(src_hw: Tuple[int, int], dst_hw: Tuple[int, int] = STREAMPETR_INPUT_HW, *, + geometry: Optional[TriangleGeometry] = None, fma: bool = False) -> TriangleResizeLUT: + """The (cached) tables for one source size. ``geometry`` overrides the StreamPETR rule + (:func:`autoware_resize_geometry`) for other nodes of the same kernel family (resized size, ROI start).""" + g = geometry or autoware_resize_geometry(src_hw[0], src_hw[1], dst_hw[0], dst_hw[1]) + + def make() -> TriangleResizeLUT: + rows, row_w = _axis_taps(g.roi_hw[0], g.roi_start[0], g.src_hw[0], g.resized_hw[0], fma) + cols, col_w = _axis_taps(g.roi_hw[1], g.roi_start[1], g.src_hw[1], g.resized_hw[1], fma) + return TriangleResizeLUT(g, rows, row_w, cols, col_w, bool(fma)) + + key = ("triangle", g.src_hw, g.resized_hw, g.roi_hw, g.roi_start, bool(fma)) + return _TRIANGLE_LUTS.get(key, make) + + +# --------------------------------------------------------------------------------------------- the preset + + +def _preset(preset: Any) -> StreamPETRPreset: + if isinstance(preset, StreamPETRPreset): + return preset + try: + return STREAMPETR_PRESETS[str(preset)] + except KeyError: + raise ValueError(f"unknown normalisation preset {preset!r}; one of {sorted(STREAMPETR_PRESETS)}") from None + + +def streampetr_preprocess(image: np.ndarray, *, preset: Any = "autoware_main", channels: str = "rgb", + dst_hw: Tuple[int, int] = STREAMPETR_INPUT_HW, fma: bool = False, + out: Optional[np.ndarray] = None) -> np.ndarray: + """One camera image -> the float32 network input ``(3, 480, 640)`` of ``autoware_camera_streampetr``. + + ``image``: uint8 ``(H, W, 3)`` in the channel order ``channels`` (``"rgb"``: :func:`ttaw.io.load_image`, an + ``rgb8`` message; ``"bgr"``: OpenCV, a ``bgr8`` message). The preset fixes the plane order the kernel sees and + the statistics (:data:`STREAMPETR_PRESETS`); ``fma`` the contraction model (module docstring).""" + p = _preset(preset) + img = np.asarray(image) + if img.ndim != 3 or img.shape[2] != 3 or img.dtype != np.uint8: + raise ValueError(f"expected an (H, W, 3) uint8 image, got {img.dtype} {img.shape}") + if channels not in ("rgb", "bgr"): + raise ValueError("channels must be 'rgb' or 'bgr'") + if channels != p.planes: + img = img[:, :, ::-1] + lut = triangle_resize_lut(img.shape[:2], dst_hw, fma=fma) + v = lut.apply(img) # (3, H, W) float32 + mean = p.mean_f32()[:, None, None] + std = p.std_f32()[:, None, None] + res = ((v - mean).astype(F32) / std).astype(F32) # (sum_c - mean[c]) / std[c] + if out is not None: + if out.shape != res.shape or out.dtype != F32: + raise ValueError(f"out must be float32 {res.shape}, got {out.dtype} {out.shape}") + out[...] = res + return out + return res + + +def streampetr_preprocess_batch(images: Sequence[np.ndarray], *, preset: Any = "autoware_main", + channels: str = "rgb", dst_hw: Tuple[int, int] = STREAMPETR_INPUT_HW, + fma: bool = False, workers: int = 1, out: Optional[np.ndarray] = None) -> np.ndarray: + """The cameras of one frame -> float32 ``(N, 3, 480, 640)`` (each camera may have its own size). ``workers`` + threads over the cameras (numpy releases the GIL in the array passes).""" + imgs = list(images) + shape = (len(imgs), 3) + tuple(int(v) for v in dst_hw) + if out is None: + out = np.empty(shape, F32) + elif out.shape != shape or out.dtype != F32: + raise ValueError(f"out must be float32 {shape}, got {out.dtype} {out.shape}") + + def one(i: int) -> None: + streampetr_preprocess(imgs[i], preset=preset, channels=channels, dst_hw=dst_hw, fma=fma, out=out[i]) + + if workers > 1 and len(imgs) > 1: + with ThreadPoolExecutor(max_workers=min(int(workers), len(imgs))) as ex: + list(ex.map(one, range(len(imgs)))) + else: + for i in range(len(imgs)): + one(i) + return out diff --git a/code/tt_diffusion_planner/ttaw/io.py b/code/tt_diffusion_planner/ttaw/io.py new file mode 100644 index 0000000000000000000000000000000000000000..b017ba7d2ab8c1fc18de0fefc0b464b685825045 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/io.py @@ -0,0 +1,877 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C08 I/O: input decoding and output encoding shared by the Python API and the HTTP server of every bundle. + +One module, two callers: the bundle's ``api`` (Python objects: paths, bytes, numpy arrays) and its ``server.app`` +(JSON envelopes with base64 payloads). Both go through the same decoders, so ``model(...)`` and ``POST /predict`` +see bit-identical inputs (BUNDLE_CONVENTIONS.md section 7). + +Rules: + +- numpy only (PIL is imported lazily for images). No ttnn, no torch, no network: this module is imported by the + image's build-time ``verify:`` step and by host tests. +- Data parsing only: ``np.load(..., allow_pickle=False)``, no ``eval`` / ``pickle``. Every client mistake raises + :class:`InputError` (the server maps it to HTTP 400). + +Point-cloud envelope (``"points"`` in the request):: + + {"format": "bin" | "npy" | "npz" | "pcd" | "list", + "data": "", # all formats except "list" + "values": [[x, y, z, intensity], ...], # "list" only (small clouds) + "fields": ["x", "y", "z", "intensity"], # column names: required meaning for "bin" / "list" / plain arrays + "dtype": "float32", # "bin" only: float32 | float64 + "key": "points", # "npz" only (default: "points", else the first array) + "frame_id": "base_link"} + +Transforms (``T__from_``, 4x4, metres) accept several spellings, see :func:`parse_transform`. +""" +from __future__ import annotations + +import base64 +import binascii +import io +import json +import math +import re +from dataclasses import dataclass, field +from pathlib import Path +from typing import Any, Iterable, Mapping, Optional, Sequence, Union + +import numpy as np + +__all__ = [ + "InputError", "PointCloud", "CameraImage", "POINT_FORMATS", "DEFAULT_POINT_FIELDS", "MAX_POINTS_DEFAULT", + "b64decode", "decode_points", "load_points", "decode_image", "load_image", "decode_cameras", "load_camera", + "load_cameras", "decode_rois", "parse_transform", + "parse_intrinsics", "resolve_calibration", "check_named_arrays", "decode_named_arrays", "load_named_arrays", + "encode_array", "encode_png", "to_jsonable", +] + +POINT_FORMATS = ("bin", "npy", "npz", "pcd", "list") +DEFAULT_POINT_FIELDS = ("x", "y", "z", "intensity") # Autoware's common LiDAR input layout +MAX_POINTS_DEFAULT = 2_000_000 # hard cap per cloud (memory guard, not a model limit) + + +class InputError(ValueError): + """A client-side input problem. The HTTP server answers 400 with ``str(err)``.""" + + +# ----------------------------------------------------------------------------------------------- base64 + +_B64_JUNK = re.compile(r"\s+") + + +def b64decode(s: Any, *, field: str = "data", max_bytes: Optional[int] = None) -> bytes: + """Decode standard or URL-safe base64; tolerates a ``data:...;base64,`` prefix, whitespace and missing padding.""" + if not isinstance(s, str) or not s: + raise InputError(f"{field} must be a non-empty base64 string") + if s.startswith("data:"): + comma = s.find(",") + if comma < 0: + raise InputError(f"{field}: malformed data: URI") + s = s[comma + 1:] + s = _B64_JUNK.sub("", s).replace("-", "+").replace("_", "/") + s += "=" * (-len(s) % 4) + if max_bytes is not None and len(s) * 3 // 4 > max_bytes + 3: + raise InputError(f"{field} is larger than {max_bytes} bytes") + try: + raw = base64.b64decode(s, validate=True) + except (binascii.Error, ValueError) as e: + raise InputError(f"{field} is not valid base64: {e}") from None + if not raw: + raise InputError(f"{field} is empty") + return raw + + +# ---------------------------------------------------------------------------------------- point clouds + +@dataclass +class PointCloud: + """An (N, C) float32 point array plus its column names and coordinate frame. Rows with a non-finite ``x``, + ``y`` or ``z`` (the first three columns when the cloud has no such names) are dropped at construction (PCL + ``is_dense=false`` clouds); the remaining rows keep their input order.""" + + points: np.ndarray + fields: tuple + frame_id: str = "base_link" + + def __post_init__(self) -> None: + self.fields = tuple(self.fields) + if self.points.ndim != 2 or self.points.shape[1] != len(self.fields): + raise InputError(f"points have shape {tuple(self.points.shape)} but {len(self.fields)} fields " + f"{list(self.fields)}") + if self.points.dtype != np.float32: + self.points = self.points.astype(np.float32) + named = [self.fields.index(n) for n in ("x", "y", "z") if n in self.fields] + xyz = self.points[:, named] if named else self.points[:, :3] + if not np.isfinite(xyz).all(): + keep = np.isfinite(xyz).all(axis=1) + self.points = np.ascontiguousarray(self.points[keep]) + + def __len__(self) -> int: + return int(self.points.shape[0]) + + def select(self, names: Sequence[str], *, fill: Optional[Mapping[str, float]] = None) -> np.ndarray: + """Columns ``names`` in that order (C-contiguous float32). A missing column is an error unless ``fill`` + gives it a constant (e.g. ``{"time_lag": 0.0}`` for a single sweep).""" + cols = [] + index = {n: i for i, n in enumerate(self.fields)} + for n in names: + if n in index: + cols.append(self.points[:, index[n]]) + elif fill is not None and n in fill: + cols.append(np.full(len(self), fill[n], np.float32)) + else: + raise InputError(f"point field {n!r} is required; the cloud has {list(self.fields)}") + return np.ascontiguousarray(np.stack(cols, axis=1), dtype=np.float32) + + +_PCD_TYPES = {("F", 4): " bytes: + """liblzf decompression (the codec of PCD ``DATA binary_compressed``). Pure Python: fine for test clouds, + slow (~1 s per 10 MB) for big ones -- send ``binary`` PCD or ``.npy`` when latency matters.""" + out = bytearray(out_len) + ip = op = 0 + n = len(data) + while ip < n: + ctrl = data[ip] + ip += 1 + if ctrl < 32: # literal run of ctrl + 1 bytes + ln = ctrl + 1 + if ip + ln > n or op + ln > out_len: + raise InputError("pcd: corrupt LZF stream (literal run)") + out[op:op + ln] = data[ip:ip + ln] + ip += ln + op += ln + else: # back reference + ln = ctrl >> 5 + ref = op - ((ctrl & 0x1F) << 8) - 1 + if ln == 7: + if ip >= n: + raise InputError("pcd: corrupt LZF stream (length)") + ln += data[ip] + ip += 1 + if ip >= n: + raise InputError("pcd: corrupt LZF stream (offset)") + ref -= data[ip] + ip += 1 + ln += 2 + if ref < 0 or op + ln > out_len: + raise InputError("pcd: corrupt LZF stream (back reference)") + if ref + ln <= op: + out[op:op + ln] = out[ref:ref + ln] + else: # overlapping copy: byte by byte, as liblzf does + for k in range(ln): + out[op + k] = out[ref + k] + op += ln + if op != out_len: + raise InputError(f"pcd: LZF produced {op} bytes, header says {out_len}") + return bytes(out) + + +def _parse_pcd(raw: bytes) -> tuple: + """PCD v0.7 (ascii / binary / binary_compressed) -> (float32 (N, C) array, field names).""" + header: dict = {} + pos = 0 + while True: + end = raw.find(b"\n", pos) + if end < 0: + raise InputError("pcd: header has no DATA line") + line = raw[pos:end].decode("ascii", "replace").strip() + pos = end + 1 + if not line or line.startswith("#"): + continue + key, _, rest = line.partition(" ") + header[key.upper()] = rest.split() + if key.upper() == "DATA": + break + try: + names = header["FIELDS"] + sizes = [int(v) for v in header["SIZE"]] + types = [v.upper() for v in header["TYPE"]] + counts = [int(v) for v in header.get("COUNT", ["1"] * len(names))] + npts = int(header.get("POINTS", [0])[0]) or int(header["WIDTH"][0]) * int(header.get("HEIGHT", ["1"])[0]) + mode = header["DATA"][0].lower() + except (KeyError, ValueError, IndexError) as e: + raise InputError(f"pcd: incomplete header ({e})") from None + if not (len(names) == len(sizes) == len(types) == len(counts)): + raise InputError("pcd: FIELDS/SIZE/TYPE/COUNT lengths differ") + cols, dt = [], [] + for n, s, t, c in zip(names, sizes, types, counts): + npt = _PCD_TYPES.get((t, s)) + if npt is None: + raise InputError(f"pcd: unsupported field type {t}{s} for {n!r}") + fname = n if n != "_" else f"_pad{len(dt)}" + dt.append((fname, npt, (c,)) if c > 1 else (fname, npt)) + if n != "_": + cols += [n] if c == 1 else [f"{n}_{k}" for k in range(c)] + dtype = np.dtype(dt) + body = raw[pos:] + if mode == "ascii": + try: + arr = np.loadtxt(io.StringIO(body.decode("ascii")), dtype=np.float64, ndmin=2) + except ValueError as e: + raise InputError(f"pcd: bad ascii body ({e})") from None + flat_names = [n if c == 1 else f"{n}_{k}" for n, c in zip(names, counts) for k in range(c)] + keep = [i for i, n in enumerate(flat_names) if not n.startswith("_")] + if arr.shape[1] != len(flat_names): + raise InputError(f"pcd: ascii rows have {arr.shape[1]} values, header declares {len(flat_names)}") + return arr[:, keep].astype(np.float32), tuple(cols) + if mode == "binary": + need = dtype.itemsize * npts + if len(body) < need: + raise InputError(f"pcd: binary body has {len(body)} bytes, needs {need}") + rec = np.frombuffer(body, dtype=dtype, count=npts) + elif mode == "binary_compressed": + if len(body) < 8: + raise InputError("pcd: binary_compressed body too short") + csize, usize = np.frombuffer(body[:8], " tuple: + cols, names = [], [] + for name in rec.dtype.names: + if name.startswith("_pad"): + continue + v = np.asarray(rec[name], dtype=np.float32) + if v.ndim == 1: + cols.append(v) + names.append(name) + else: + for k in range(v.shape[1]): + cols.append(v[:, k]) + names.append(f"{name}_{k}") + return np.stack(cols, axis=1) if cols else np.zeros((len(rec), 0), np.float32), tuple(names) + + +def _array_to_cloud(arr: np.ndarray, fields: Optional[Sequence[str]], frame_id: str, + default_fields: Sequence[str]) -> PointCloud: + if arr.dtype.names: # structured array (npy / npz written from a ROS PointCloud2 or PCL) + pts, names = _structured_to_2d(arr.reshape(-1)) + return PointCloud(pts, names, frame_id) + arr = np.asarray(arr) + if arr.ndim != 2: + raise InputError(f"points must be a 2-D (N, C) array, got shape {tuple(arr.shape)}") + if not np.issubdtype(arr.dtype, np.number): + raise InputError(f"points must be numeric, got dtype {arr.dtype}") + names = tuple(fields) if fields else tuple(default_fields[: arr.shape[1]]) + if len(names) != arr.shape[1]: + raise InputError(f"points have {arr.shape[1]} columns; give 'fields' (default {list(default_fields)})") + return PointCloud(np.ascontiguousarray(arr, dtype=np.float32), names, frame_id) + + +def _points_from_bytes(raw: bytes, fmt: str, *, fields: Optional[Sequence[str]] = None, dtype: str = "float32", + key: Optional[str] = None, frame_id: str = "base_link", + default_fields: Sequence[str] = DEFAULT_POINT_FIELDS) -> PointCloud: + if fmt == "bin": # raw little-endian rows, KITTI / Autoware dump style: N x len(fields) + names = tuple(fields) if fields else tuple(default_fields) + if dtype not in ("float32", "float64"): + raise InputError("bin dtype must be float32 or float64") + item = 4 if dtype == "float32" else 8 + row = item * len(names) + if len(raw) % row: + raise InputError(f"bin payload of {len(raw)} bytes is not a multiple of {len(names)} x {dtype} " + f"({row} bytes per point); check 'fields' (got {list(names)})") + arr = np.frombuffer(raw, dtype=" PointCloud: + """The JSON point-cloud envelope (module docstring) -> :class:`PointCloud`.""" + if not isinstance(spec, Mapping): + raise InputError("points must be an object {format, data, fields, ...}") + fmt = str(spec.get("format") or "bin").lower() + frame_id = str(spec.get("frame_id") or "base_link") + fields = spec.get("fields") + if fields is not None and (not isinstance(fields, (list, tuple)) or not all(isinstance(f, str) for f in fields)): + raise InputError("points.fields must be a list of column names") + if fmt == "list": + values = spec.get("values") + if not isinstance(values, list) or not values: + raise InputError("points.values must be a non-empty list of rows for format 'list'") + try: + arr = np.asarray(values, dtype=np.float32) + except (TypeError, ValueError) as e: + raise InputError(f"points.values: {e}") from None + cloud = _array_to_cloud(arr, fields, frame_id, default_fields) + else: + raw = b64decode(spec.get("data"), field="points.data", max_bytes=max_bytes) + cloud = _points_from_bytes(raw, fmt, fields=fields, dtype=str(spec.get("dtype") or "float32"), + key=spec.get("key"), frame_id=frame_id, default_fields=default_fields) + if len(cloud) > max_points: + raise InputError(f"{len(cloud)} points exceed the limit of {max_points}") + if len(cloud) == 0: + raise InputError("the point cloud is empty") + return cloud + + +PointsLike = Union[str, Path, bytes, np.ndarray, Mapping[str, Any], PointCloud] + + +def load_points(source: PointsLike, *, fields: Optional[Sequence[str]] = None, fmt: Optional[str] = None, + frame_id: str = "base_link", default_fields: Sequence[str] = DEFAULT_POINT_FIELDS) -> PointCloud: + """Python-API input: a path (.bin/.npy/.npz/.pcd), raw file bytes (give ``fmt``), an (N, C) array (or torch + tensor), a structured array, the JSON envelope, or a :class:`PointCloud`.""" + if isinstance(source, PointCloud): + return source + if isinstance(source, Mapping): + return decode_points(source, default_fields=default_fields) + if isinstance(source, (str, Path)): + p = Path(source) + ext = (fmt or p.suffix.lstrip(".")).lower() + if p.name.endswith(".pcd.bin"): # nuScenes naming: raw float32 x5 + ext = "bin" + return _points_from_bytes(p.read_bytes(), ext, fields=fields, frame_id=frame_id, default_fields=default_fields) + if isinstance(source, (bytes, bytearray, memoryview)): + if not fmt: + raise InputError("raw bytes need fmt= ('bin' | 'npy' | 'npz' | 'pcd')") + return _points_from_bytes(bytes(source), fmt, fields=fields, frame_id=frame_id, default_fields=default_fields) + if hasattr(source, "detach") and hasattr(source, "cpu"): # torch tensor, without importing torch + source = source.detach().cpu().numpy() + return _array_to_cloud(np.asarray(source), fields, frame_id, default_fields) + + +# ---------------------------------------------------------------------------------------------- images + +def _pil_to_rgb(im) -> np.ndarray: + if im.mode != "RGB": + im = im.convert("RGB") + return np.asarray(im, dtype=np.uint8) + + +def decode_image(raw: bytes, *, fmt: str = "auto", max_side: int = 8192) -> np.ndarray: + """PNG / JPEG (or .npy uint8 HxWx3) bytes -> RGB uint8 (H, W, 3).""" + if fmt == "npy": + try: + arr = np.load(io.BytesIO(raw), allow_pickle=False) + except ValueError as e: + raise InputError(f"image npy: {e}") from None + return load_image(arr) + from PIL import Image, UnidentifiedImageError # lazy: the API may never see an image + + try: + with Image.open(io.BytesIO(raw)) as im: + if max(im.size) > max_side: + raise InputError(f"image size {im.size} exceeds max side {max_side}") + return _pil_to_rgb(im) + except (UnidentifiedImageError, OSError) as e: + raise InputError(f"image could not be decoded as PNG/JPEG: {e}") from None + + +def load_image(image: Any) -> np.ndarray: + """Python-API image input: path, bytes, PIL image, uint8 HxWx3 / 3xHxW array or tensor, float array in [0, 1] + -> RGB uint8 (H, W, 3).""" + if isinstance(image, (str, Path)): + return decode_image(Path(image).read_bytes()) + if isinstance(image, (bytes, bytearray, memoryview)): + return decode_image(bytes(image)) + if hasattr(image, "mode") and hasattr(image, "size") and hasattr(image, "convert"): # PIL + return _pil_to_rgb(image) + if hasattr(image, "detach") and hasattr(image, "cpu"): + image = image.detach().cpu().numpy() + a = np.asarray(image) + if a.ndim == 3 and a.shape[0] in (1, 3) and a.shape[2] not in (1, 3): + a = np.transpose(a, (1, 2, 0)) + if a.ndim == 2: + a = np.repeat(a[:, :, None], 3, axis=2) + if a.ndim != 3 or a.shape[2] not in (1, 3, 4): + raise InputError(f"image array must be HxWx3 (or 3xHxW / HxW), got {a.shape}") + a = a[:, :, :3] if a.shape[2] != 1 else np.repeat(a, 3, axis=2) + if np.issubdtype(a.dtype, np.floating): + a = np.clip(np.round(a * 255.0), 0, 255) + return np.ascontiguousarray(a, dtype=np.uint8) + + +@dataclass +class CameraImage: + """One camera of a (multi-)camera request: pixels + calibration, in the order the model expects.""" + + name: str + image: np.ndarray # RGB uint8 (H, W, 3) + intrinsics: Optional[np.ndarray] = None # (3, 3) float64, pixels + T_ref_from_camera: Optional[np.ndarray] = None # (4, 4) float64: camera -> reference frame (base_link / lidar) + distortion: Optional[dict] = None + timestamp_s: Optional[float] = None + extra: dict = field(default_factory=dict) + + +def decode_cameras(images: Iterable[Mapping[str, Any]], calibration: Optional[Mapping[str, Any]] = None, *, + order: Optional[Sequence[str]] = None, require_calibration: bool = True, + max_bytes: Optional[int] = None) -> list: + """``images[]`` (+ optional ``calibration.cameras``) -> list of :class:`CameraImage` in ``order``. + + Each image entry: ``{"camera": "CAM_FRONT", "data": , "format": "auto|png|jpeg|npy", + "intrinsics": ..., "T_ref_from_camera": ..., "distortion": {...}, "timestamp_s": ...}``. Inline calibration + wins over ``calibration.cameras[]``. ``order`` (the model's fixed camera order) reorders and checks + that every camera is present exactly once.""" + calib = dict((calibration or {}).get("cameras") or {}) + out = [] + for i, entry in enumerate(images or []): + if not isinstance(entry, Mapping): + raise InputError(f"images[{i}] must be an object") + name = str(entry.get("camera") or (order[i] if order and i < len(order) else f"cam{i}")) + raw = b64decode(entry.get("data"), field=f"images[{i}].data", max_bytes=max_bytes) + cam_cal = {**dict(calib.get(name) or {}), + **{k: v for k, v in entry.items() if k not in ("camera", "data", "format")}} + k = cam_cal.get("intrinsics") + t = cam_cal.get("T_ref_from_camera", cam_cal.get("extrinsics")) + if require_calibration and (k is None or t is None): + raise InputError(f"camera {name!r} needs 'intrinsics' and 'T_ref_from_camera' (inline or in calibration)") + out.append(CameraImage( + name=name, image=decode_image(raw, fmt=str(entry.get("format") or "auto")), + intrinsics=None if k is None else parse_intrinsics(k, field=f"{name}.intrinsics"), + T_ref_from_camera=None if t is None else parse_transform(t, field=f"{name}.T_ref_from_camera"), + distortion=cam_cal.get("distortion"), timestamp_s=cam_cal.get("timestamp_s"))) + if order: + by_name = {c.name: c for c in out} + if len(by_name) != len(out): + raise InputError("duplicate camera names in images[]") + missing = [n for n in order if n not in by_name] + extra = [n for n in by_name if n not in order] + if missing or extra: + raise InputError(f"cameras must be exactly {list(order)}; missing {missing}, unexpected {extra}") + out = [by_name[n] for n in order] + return out + + +_CAMERA_KEYS = ("intrinsics", "T_ref_from_camera", "extrinsics", "distortion", "timestamp_s") + + +def _with_calibration(cam: CameraImage, cal: Mapping[str, Any], require: bool) -> CameraImage: + """``cam`` with missing intrinsics / extrinsics / distortion / timestamp taken from ``cal`` (its calibration + entry), parsed and checked; ``require`` refuses a camera without K and T_ref_from_camera.""" + k = cam.intrinsics if cam.intrinsics is not None else cal.get("intrinsics") + t = cam.T_ref_from_camera if cam.T_ref_from_camera is not None else cal.get("T_ref_from_camera", + cal.get("extrinsics")) + if require and (k is None or t is None): + raise InputError(f"camera {cam.name!r} needs 'intrinsics' and 'T_ref_from_camera' (inline or in calibration)") + return CameraImage( + name=cam.name, image=cam.image, + intrinsics=None if k is None else parse_intrinsics(k, field=f"{cam.name}.intrinsics"), + T_ref_from_camera=None if t is None else parse_transform(t, field=f"{cam.name}.T_ref_from_camera"), + distortion=cam.distortion if cam.distortion is not None else cal.get("distortion"), + timestamp_s=cam.timestamp_s if cam.timestamp_s is not None else cal.get("timestamp_s"), extra=dict(cam.extra)) + + +def load_camera(source: Any, *, name: Optional[str] = None, calibration: Optional[Mapping[str, Any]] = None, + require_calibration: bool = False, max_bytes: Optional[int] = None) -> CameraImage: + """One camera of the Python API -> :class:`CameraImage` (RGB uint8 HxWx3, K / T parsed). ``source``: a + :class:`CameraImage` (what the server decodes ``/predict`` into), a mapping ``{"camera", "image" | "path", + "intrinsics", "T_ref_from_camera", "distortion", "timestamp_s"}`` or the ``/predict`` spelling ``{"camera", + "data": , ...}``, or any image :func:`load_image` accepts (then ``name`` names it). Calibration missing + inline comes from ``calibration["cameras"][]`` (inline wins, as in :func:`decode_cameras`).""" + cams = dict((calibration or {}).get("cameras") or {}) + if isinstance(source, CameraImage): + cam = source + elif isinstance(source, Mapping): + cam_name = str(source.get("camera") or name or "") + if not cam_name: + raise InputError("every camera needs a 'camera' name") + if "data" in source: + raw = b64decode(source.get("data"), field=f"{cam_name}.data", max_bytes=max_bytes) + image = decode_image(raw, fmt=str(source.get("format") or "auto")) + else: + src = source.get("image", source.get("path")) + if src is None: + raise InputError(f"camera {cam_name!r}: no 'image', 'path' or 'data'") + image = load_image(src) + unknown = sorted(set(source) - {"camera", "data", "format", "image", "path", *_CAMERA_KEYS}) + if unknown: + raise InputError(f"camera {cam_name!r}: unknown keys {unknown}") + cam = CameraImage(name=cam_name, image=image, intrinsics=source.get("intrinsics"), + T_ref_from_camera=source.get("T_ref_from_camera", source.get("extrinsics")), + distortion=source.get("distortion"), timestamp_s=source.get("timestamp_s")) + else: + if not name: + raise InputError("an image without a camera name: pass {'camera': name, 'image': ...} or a " + "{name: image} mapping") + cam = CameraImage(name=str(name), image=load_image(source)) + return _with_calibration(cam, dict(cams.get(cam.name) or {}), require_calibration) + + +def load_cameras(images: Any, calibration: Optional[Mapping[str, Any]] = None, *, + order: Optional[Sequence[str]] = None, require_calibration: bool = True, + calib_dir: Optional[Path] = None, max_bytes: Optional[int] = None) -> list: + """The Python-API twin of :func:`decode_cameras`: ``images`` (a list of :func:`load_camera` sources, or a + ``{camera name: image}`` mapping) + ``calibration`` (``{"cameras": {name: {...}}}``, or ``{"preset": name}`` + resolved in ``calib_dir``) -> :class:`CameraImage` list, reordered to ``order`` (every camera exactly once) when + given, else in input order. ``model(images=..., calibration=...)`` and ``POST /predict`` then see the same + cameras: the server passes the CameraImages it decoded, which come back unchanged.""" + if images is None: + raise InputError("this model needs 'images'" + (f": the cameras {list(order)}" if order else "")) + calib = resolve_calibration(calibration, calib_dir) if calibration is not None else None + if isinstance(images, Mapping) and not ({"camera", "image", "path", "data"} & set(images)): + entries = [(str(k), v) for k, v in images.items()] + else: + seq = list(images) if isinstance(images, (list, tuple)) else [images] + entries = [(None, v) for v in seq] + out = [load_camera(v, name=k, calibration=calib, require_calibration=require_calibration, max_bytes=max_bytes) + for k, v in entries] + names = [c.name for c in out] + if len(set(names)) != len(names): + raise InputError(f"duplicate camera names in images: {names}") + if order: + missing = [n for n in order if n not in names] + extra = [n for n in names if n not in order] + if missing or extra: + raise InputError(f"cameras must be exactly {list(order)}; missing {missing}, unexpected {extra}") + by_name = {c.name: c for c in out} + out = [by_name[n] for n in order] + return out + + +def decode_rois(rois: Any, *, cameras: Optional[Sequence[str]] = None, labels: Optional[Sequence[str]] = None, + max_rois: int = 4096) -> list: + """2-D detections given as an input (PointPainting's ``rois``, BUNDLE_CONVENTIONS.md 7.2): + ``[{"camera", "label", "score", "box_xyxy"}]`` -> the same list checked and normalised: ``camera`` a name (one of + ``cameras`` when given), ``label`` a name of ``labels`` or an index into it (returned as both ``label`` and + ``label_id`` when ``labels`` is given), ``score`` in [0, 1] (default 1.0), ``box_xyxy`` 4 finite pixels with + x0 <= x1 and y0 <= y1. ``None`` -> ``[]``.""" + if rois is None: + return [] + if not isinstance(rois, (list, tuple)): + raise InputError("rois must be a list of {camera, label, score, box_xyxy}") + if len(rois) > max_rois: + raise InputError(f"{len(rois)} rois exceed the limit of {max_rois}") + out = [] + for i, r in enumerate(rois): + where = f"rois[{i}]" + if not isinstance(r, Mapping): + raise InputError(f"{where} must be an object") + unknown = sorted(set(r) - {"camera", "label", "label_id", "score", "box_xyxy"}) + if unknown: + raise InputError(f"{where}: unknown keys {unknown}") + cam = r.get("camera") + if not isinstance(cam, str) or not cam: + raise InputError(f"{where}.camera must be a camera name") + if cameras is not None and cam not in cameras: + raise InputError(f"{where}.camera {cam!r} is not one of {list(cameras)}") + label = r.get("label", r.get("label_id")) + if isinstance(label, bool) or not isinstance(label, (str, int)): + raise InputError(f"{where}.label must be a class name or index") + entry: dict = {"camera": cam} + if labels is not None: + names = list(labels) + if isinstance(label, int): + if not 0 <= label < len(names): + raise InputError(f"{where}.label {label} is not an index of {names}") + entry.update(label=names[label], label_id=label) + elif label in names: + entry.update(label=label, label_id=names.index(label)) + else: + raise InputError(f"{where}.label {label!r} is not one of {names}") + else: + entry["label"] = label + score = r.get("score", 1.0) + try: + score = float(score) + except (TypeError, ValueError): + raise InputError(f"{where}.score must be a number") from None + if isinstance(r.get("score"), bool) or not 0.0 <= score <= 1.0: + raise InputError(f"{where}.score must be in [0, 1]") + try: + box = [float(v) for v in r.get("box_xyxy")] + except (TypeError, ValueError): + raise InputError(f"{where}.box_xyxy must be 4 numbers") from None + if len(box) != 4 or not all(math.isfinite(v) for v in box) or box[0] > box[2] or box[1] > box[3]: + raise InputError(f"{where}.box_xyxy must be 4 finite pixels with x0 <= x1 and y0 <= y1") + entry.update(score=score, box_xyxy=box) + out.append(entry) + return out + + +# ----------------------------------------------------------------------------------------- calibration + +def _rot_from_quat_wxyz(w: float, x: float, y: float, z: float) -> np.ndarray: + n = math.sqrt(w * w + x * x + y * y + z * z) + if n < 1e-9: + raise InputError("zero-length quaternion") + w, x, y, z = w / n, x / n, y / n, z / n + return np.array([[1 - 2 * (y * y + z * z), 2 * (x * y - z * w), 2 * (x * z + y * w)], + [2 * (x * y + z * w), 1 - 2 * (x * x + z * z), 2 * (y * z - x * w)], + [2 * (x * z - y * w), 2 * (y * z + x * w), 1 - 2 * (x * x + y * y)]], dtype=np.float64) + + +def _rot_from_rpy(roll: float, pitch: float, yaw: float) -> np.ndarray: + """tf2 ``setRPY`` convention used by Autoware's sensor_kit_calibration.yaml: R = Rz(yaw) Ry(pitch) Rx(roll).""" + cr, sr, cp, sp, cy, sy = (math.cos(roll), math.sin(roll), math.cos(pitch), math.sin(pitch), math.cos(yaw), + math.sin(yaw)) + return np.array([[cy * cp, cy * sp * sr - sy * cr, cy * sp * cr + sy * sr], + [sy * cp, sy * sp * sr + cy * cr, sy * sp * cr - cy * sr], + [-sp, cp * sr, cp * cr]], dtype=np.float64) + + +def parse_transform(obj: Any, *, field: str = "transform") -> np.ndarray: + """A rigid transform -> (4, 4) float64. Accepted spellings: + + - a 4x4 (or 3x4) nested list / array, row-major, translation in the last column; + - ``{"matrix": 4x4}``; + - ``{"translation": [x, y, z], "rotation_wxyz": [w, x, y, z]}`` (nuScenes / BEVDet sample yaml), or + ``rotation_xyzw``, or ``rotation`` (3x3); + - ``{"x", "y", "z", "roll", "pitch", "yaw"}`` in metres / radians (Autoware ``sensor_kit_calibration.yaml``). + """ + try: + if isinstance(obj, Mapping): + if "matrix" in obj: + return parse_transform(obj["matrix"], field=field) + if "translation" in obj: + t = np.asarray(obj["translation"], dtype=np.float64).reshape(3) + if "rotation_wxyz" in obj: + r = _rot_from_quat_wxyz(*np.asarray(obj["rotation_wxyz"], dtype=np.float64).reshape(4)) + elif "rotation_xyzw" in obj: + x, y, z, w = np.asarray(obj["rotation_xyzw"], dtype=np.float64).reshape(4) + r = _rot_from_quat_wxyz(w, x, y, z) + elif "rotation" in obj: + r = np.asarray(obj["rotation"], dtype=np.float64).reshape(3, 3) + else: + raise InputError(f"{field}: translation needs rotation_wxyz | rotation_xyzw | rotation (3x3)") + elif {"x", "y", "z", "roll", "pitch", "yaw"} <= set(obj): + t = np.array([obj["x"], obj["y"], obj["z"]], dtype=np.float64) + r = _rot_from_rpy(float(obj["roll"]), float(obj["pitch"]), float(obj["yaw"])) + else: + raise InputError(f"{field}: unrecognised transform keys {sorted(obj)}") + m = np.eye(4) + m[:3, :3], m[:3, 3] = r, t + else: + m = np.asarray(obj, dtype=np.float64) + if m.shape == (3, 4): + m = np.vstack([m, [0.0, 0.0, 0.0, 1.0]]) + if m.shape != (4, 4): + raise InputError(f"{field}: expected a 4x4 matrix, got shape {m.shape}") + except (TypeError, ValueError) as e: + if isinstance(e, InputError): + raise + raise InputError(f"{field}: {e}") from None + if not np.isfinite(m).all() or not np.allclose(m[3], [0, 0, 0, 1], atol=1e-6): + raise InputError(f"{field}: last row must be [0, 0, 0, 1]") + r = m[:3, :3] + if not np.allclose(r.T @ r, np.eye(3), atol=1e-3) or np.linalg.det(r) < 0: + raise InputError(f"{field}: rotation part is not a proper rotation") + return m + + +def parse_intrinsics(obj: Any, *, field: str = "intrinsics") -> np.ndarray: + """3x3 K, a 3x4 projection P (its left 3x3), or ``{"fx", "fy", "cx", "cy"[, "skew"]}`` -> (3, 3) float64.""" + try: + if isinstance(obj, Mapping): + k = np.array([[obj["fx"], obj.get("skew", 0.0), obj["cx"]], [0.0, obj["fy"], obj["cy"]], [0.0, 0.0, 1.0]], + dtype=np.float64) + else: + k = np.asarray(obj, dtype=np.float64) + if k.shape == (3, 4): + k = k[:, :3] + if k.shape != (3, 3): + raise InputError(f"{field}: expected 3x3, got shape {k.shape}") + except (KeyError, TypeError, ValueError) as e: + if isinstance(e, InputError): + raise + raise InputError(f"{field}: {e}") from None + if not np.isfinite(k).all() or k[0, 0] <= 0 or k[1, 1] <= 0: + raise InputError(f"{field}: fx and fy must be positive") + return k + + +_PRESET_NAME = re.compile(r"[A-Za-z0-9_.-]+") + + +def resolve_calibration(calibration: Optional[Mapping[str, Any]], calib_dir: Optional[Path]) -> Optional[dict]: + """``{"preset": "", ...}`` -> the preset file ``/.json`` merged with the other keys + (request keys win); any other mapping is returned as a dict; ``None`` stays ``None``.""" + if calibration is None: + return None + if "preset" not in calibration: + return dict(calibration) + name = str(calibration["preset"]) + if calib_dir is None or not _PRESET_NAME.fullmatch(name) or not (Path(calib_dir) / f"{name}.json").is_file(): + raise InputError(f"unknown calibration preset {name!r}; see /info input.calibration_presets") + base = json.loads((Path(calib_dir) / f"{name}.json").read_text()) + return {**base, **{k: v for k, v in calibration.items() if k != "preset"}} + + +# ------------------------------------------------------------------------------ named tensors (planner) + +def check_named_arrays(arrays: Mapping[str, Any], schema: Mapping[str, tuple]) -> dict: + """``{name: array}`` checked against ``schema`` (name -> (shape with ``None`` for free dims, dtype)): every listed + name is required, unlisted names are rejected, shapes must match and values must convert to the dtype without + loss of kind (numbers only; a float into an integer slot must be integral, values must be finite). Returns + C-contiguous arrays of the schema dtypes; every problem raises :class:`InputError`.""" + if not isinstance(arrays, Mapping): + raise InputError("inputs must be an object {name: array}") + unknown = sorted(set(arrays) - set(schema)) + missing = sorted(set(schema) - set(arrays)) + if unknown or missing: + raise InputError(f"inputs: missing {missing}, unexpected {unknown}") + out = {} + for name, (shape, dtype) in schema.items(): + try: + a = np.asarray(arrays[name]) + except (TypeError, ValueError) as e: + raise InputError(f"inputs[{name!r}]: {e}") from None + if len(a.shape) != len(shape) or any(s is not None and s != d for s, d in zip(shape, a.shape)): + raise InputError(f"inputs[{name!r}] has shape {tuple(a.shape)}, expected {tuple(shape)}") + if not (np.issubdtype(a.dtype, np.number) or a.dtype == np.bool_): + raise InputError(f"inputs[{name!r}] must be numeric, got dtype {a.dtype}") + target = np.dtype(dtype) + if a.size and np.issubdtype(a.dtype, np.inexact) and not np.isfinite(a).all(): + raise InputError(f"inputs[{name!r}] holds non-finite values") + if a.size and np.issubdtype(target, np.integer) and a.dtype != np.bool_: + if np.issubdtype(a.dtype, np.inexact) and not np.array_equal(a, np.round(a)): + raise InputError(f"inputs[{name!r}] must hold integers ({target})") + info = np.iinfo(target) + if a.min() < info.min or a.max() > info.max: + raise InputError(f"inputs[{name!r}] has values outside the {target} range") + out[name] = np.ascontiguousarray(a.astype(target, copy=False)) + return out + + +def decode_named_arrays(spec: Mapping[str, Any], schema: Optional[Mapping[str, tuple]] = None, *, + max_bytes: Optional[int] = None) -> dict: + """``{"format": "npz", "data": }`` or ``{"format": "json", "arrays": {name: nested list}}`` -> + ``{name: ndarray}``. With ``schema`` the arrays are checked and cast by :func:`check_named_arrays`.""" + if not isinstance(spec, Mapping): + raise InputError("inputs must be an object {format, data | arrays}") + fmt = str(spec.get("format") or "npz").lower() + if fmt == "npz": + raw = b64decode(spec.get("data"), field="inputs.data", max_bytes=max_bytes) + try: + z = np.load(io.BytesIO(raw), allow_pickle=False) + arrays = {k: z[k] for k in z.files} + except (ValueError, OSError) as e: + raise InputError(f"inputs npz: {e}") from None + elif fmt == "json": + src = spec.get("arrays") + if not isinstance(src, Mapping): + raise InputError("inputs.arrays must be an object {name: nested list}") + arrays = {} + for k, v in src.items(): + try: + arrays[k] = np.asarray(v) + except (TypeError, ValueError) as e: + raise InputError(f"inputs.arrays[{k!r}]: {e}") from None + else: + raise InputError(f"inputs.format {fmt!r} must be 'npz' or 'json'") + return arrays if schema is None else check_named_arrays(arrays, schema) + + +def load_named_arrays(source: Any, schema: Optional[Mapping[str, tuple]] = None) -> dict: + """Python-API named inputs (planner): a ``{name: array}`` mapping, the JSON envelope of + :func:`decode_named_arrays`, an ``.npz`` path or its bytes -> ``{name: ndarray}`` (checked against ``schema`` + when given, so ``model(inputs=...)`` and ``POST /predict`` accept and refuse the same inputs).""" + if isinstance(source, Mapping) and "format" in source and ("data" in source or "arrays" in source): + return decode_named_arrays(source, schema) + if isinstance(source, (str, Path, bytes, bytearray, memoryview)): + try: + payload = Path(source).read_bytes() if isinstance(source, (str, Path)) else bytes(source) + with np.load(io.BytesIO(payload), allow_pickle=False) as z: + arrays = {k: z[k] for k in z.files} + except (ValueError, OSError) as e: + raise InputError(f"inputs npz: {e}") from None + elif isinstance(source, Mapping): + arrays = {str(k): (v.detach().cpu().numpy() if hasattr(v, "detach") else v) for k, v in source.items()} + else: + raise InputError(f"inputs must be a mapping, an .npz path or bytes, not {type(source).__name__}") + return arrays if schema is None else check_named_arrays(arrays, schema) + + +# ------------------------------------------------------------------------------------------- outputs + +def encode_array(arr: np.ndarray, *, fmt: str = "npz", key: str = "array") -> dict: + """A dense output for the JSON envelope: ``{"format", "key", "dtype", "shape", "data"}``. + + ``npz`` (default): base64 of ``np.savez`` -> ``np.load(io.BytesIO(base64.b64decode(d["data"])))[d["key"]]``; + ``npz_compressed``: the same with ``savez_compressed`` (smaller, slower); ``raw``: base64 of the little-endian + bytes (``np.frombuffer(..., dtype).reshape(shape)``); ``list``: nested JSON lists (small arrays only).""" + arr = np.asarray(arr) + head = {"format": fmt, "key": key, "dtype": str(arr.dtype), "shape": list(arr.shape)} + if fmt in ("npz", "npz_compressed"): + buf = io.BytesIO() + (np.savez_compressed if fmt == "npz_compressed" else np.savez)(buf, **{key: arr}) + return {**head, "data": base64.b64encode(buf.getvalue()).decode("ascii")} + if fmt == "raw": + a = arr.astype(arr.dtype.newbyteorder("<"), copy=False) + return {**head, "data": base64.b64encode(np.ascontiguousarray(a).tobytes()).decode("ascii")} + if fmt == "list": + return {**head, "data": arr.tolist()} + raise ValueError(f"unknown array encoding {fmt!r}") + + +def encode_png(image: np.ndarray, *, key: str = "image") -> dict: + """A uint8 mask / label map (H, W) or RGB image (H, W, 3) as base64 PNG (lossless): + ``{"format": "png", "key", "dtype", "shape", "data"}``.""" + from PIL import Image + + a = np.asarray(image) + if a.dtype != np.uint8: + if a.size and (a.min() < 0 or a.max() > 255): + raise ValueError("encode_png needs values in 0..255 (use encode_array for wider label ranges)") + a = a.astype(np.uint8) + if a.ndim not in (2, 3): + raise ValueError(f"encode_png expects (H, W) or (H, W, 3), got {a.shape}") + buf = io.BytesIO() + Image.fromarray(a).save(buf, format="PNG") + return {"format": "png", "key": key, "dtype": "uint8", "shape": list(a.shape), + "data": base64.b64encode(buf.getvalue()).decode("ascii")} + + +def to_jsonable(obj: Any) -> Any: + """numpy scalars / arrays -> plain Python, recursively. float32 / float16 scalars are rounded to 6 significant + digits (their precision); float64 scalars and array elements are kept exactly (timestamps, poses).""" + if isinstance(obj, Mapping): + return {str(k): to_jsonable(v) for k, v in obj.items()} + if isinstance(obj, (list, tuple)): + return [to_jsonable(v) for v in obj] + if isinstance(obj, np.ndarray): + return to_jsonable(obj.tolist()) + if isinstance(obj, (np.float32, np.float16)): + return float(f"{float(obj):.6g}") + if isinstance(obj, np.floating): + return float(obj) + if isinstance(obj, np.integer): + return int(obj) + if isinstance(obj, np.bool_): + return bool(obj) + if isinstance(obj, Path): + return str(obj) + return obj diff --git a/code/tt_diffusion_planner/ttaw/knobs.py b/code/tt_diffusion_planner/ttaw/knobs.py new file mode 100644 index 0000000000000000000000000000000000000000..beb8c02586bbca51f9ca92bc16f6150348740e14 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/knobs.py @@ -0,0 +1,177 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Optimization knobs: declared once, read from the environment once at model build, each an A/B switch. + +The rule every port follows (PLAN.md section 1.1; RP section 1.2): a knob's default is the measured best, the +server pins the full set in ``tt-model.yaml serve.env`` (:meth:`Knobs.serve_env` renders it), and a host test can +check the pins against the Python defaults. Setting an ``experiment`` knob logs a warning, so an A/B run is never +mistaken for the published configuration. + +Example:: + + KNOBS = Knobs("CENTERPOINT", [ + Knob("FUSED_HEAD", True, "merged 64->384->15 head convs (False = one conv per head)"), + Knob("BFP8_WEIGHTS", False, "bfp8 backbone weights", experiment=True), + Knob("NUM_CQS", 1, "command queues", choices=(1, 2)), + ]) + knobs = KNOBS.read() # at build; KnobValues is immutable + if knobs.FUSED_HEAD: ... +""" +from __future__ import annotations + +import logging +import os +from dataclasses import dataclass +from typing import Any, Dict, Iterable, Iterator, Mapping, Optional, Tuple + +__all__ = ["Knob", "Knobs", "KnobValues", "parse_bool"] + +log = logging.getLogger(__name__) + +_TRUE = {"1", "true", "yes", "on", "y"} +_FALSE = {"0", "false", "no", "off", "n", ""} + + +def parse_bool(text: str) -> bool: + value = str(text).strip().lower() + if value in _TRUE: + return True + if value in _FALSE: + return False + raise ValueError(f"{text!r} is not a boolean (1/0, true/false, on/off, yes/no)") + + +@dataclass(frozen=True) +class Knob: + """One switch. ``name`` is the suffix of the env variable ``_``; the type is the default's.""" + + name: str + default: Any + doc: str = "" + choices: Optional[Tuple[Any, ...]] = None + experiment: bool = False + + def __post_init__(self) -> None: + if not self.name or not self.name.replace("_", "").isalnum() or self.name.upper() != self.name: + raise ValueError(f"knob name {self.name!r} must be UPPER_SNAKE_CASE") + if self.choices is not None and self.default not in self.choices: + raise ValueError(f"knob {self.name}: default {self.default!r} not in {self.choices}") + + def parse(self, text: str) -> Any: + kind = type(self.default) + try: + if kind is bool: + value: Any = parse_bool(text) + elif kind is int: + value = int(text.strip()) + elif kind is float: + value = float(text.strip()) + else: + value = text.strip() + except ValueError as exc: + raise ValueError(f"{self.name}={text!r}: {exc}") from None + if self.choices is not None and value not in self.choices: + raise ValueError(f"{self.name}={value!r}: expected one of {self.choices}") + return value + + def render(self, value: Any) -> str: + if isinstance(value, bool): + return "1" if value else "0" + return str(value) + + +class KnobValues(Mapping): + """Immutable knob values with their source (``"default"`` / ``"env"`` / ``"override"``); attribute or item + access.""" + + def __init__(self, prefix: str, values: Dict[str, Any], sources: Dict[str, str]): + object.__setattr__(self, "_prefix", prefix) + object.__setattr__(self, "_values", dict(values)) + object.__setattr__(self, "_sources", dict(sources)) + + def __getattr__(self, name: str) -> Any: + if name.startswith("_"): + raise AttributeError(name) + try: + return self._values[name] + except KeyError: + raise AttributeError(f"no knob {name!r} (prefix {self._prefix})") from None + + def __setattr__(self, name: str, value: Any) -> None: + raise AttributeError("KnobValues is read-only (knobs are read once at build)") + + def __getitem__(self, name: str) -> Any: + return self._values[name] + + def __iter__(self) -> Iterator[str]: + return iter(self._values) + + def __len__(self) -> int: + return len(self._values) + + def source(self, name: str) -> str: + return self._sources[name] + + def overridden(self) -> Dict[str, Any]: + """Knobs not at their default source (set from the environment or by an explicit override).""" + return {k: v for k, v in self._values.items() if self._sources[k] != "default"} + + def as_dict(self) -> Dict[str, Any]: + return dict(self._values) + + def __repr__(self) -> str: + return f"KnobValues({self._prefix}: {self._values})" + + +class Knobs: + """The declared knob set of one model (env prefix = the bundle's ````).""" + + def __init__(self, prefix: str, knobs: Iterable[Knob]): + self.prefix = prefix.strip().upper() + self.knobs: Dict[str, Knob] = {} + for knob in knobs: + if knob.name in self.knobs: + raise ValueError(f"duplicate knob {knob.name}") + self.knobs[knob.name] = knob + + def env_name(self, name: str) -> str: + return f"{self.prefix}_{name}" + + def defaults(self) -> KnobValues: + return KnobValues(self.prefix, {n: k.default for n, k in self.knobs.items()}, + {n: "default" for n in self.knobs}) + + def read(self, env: Optional[Mapping[str, str]] = None, **overrides: Any) -> KnobValues: + """Defaults, then ``_`` environment values, then explicit ``overrides`` (e.g. from + ``from_pretrained`` keyword arguments). Raises on unparsable values; warns on experiment knobs.""" + env = os.environ if env is None else env + values, sources = {}, {} + for name, knob in self.knobs.items(): + raw = env.get(self.env_name(name)) + if raw is not None and raw.strip() != "": + values[name], sources[name] = knob.parse(raw), "env" + else: + values[name], sources[name] = knob.default, "default" + for name, value in overrides.items(): + if name not in self.knobs: + raise KeyError(f"unknown knob {name!r}; have {sorted(self.knobs)}") + knob = self.knobs[name] + values[name], sources[name] = knob.parse(knob.render(value)), "override" + experiments = [n for n, k in self.knobs.items() if k.experiment and values[n] != k.default] + if experiments: + log.warning("%s: experiment knobs set (%s); results are not the published configuration", self.prefix, + ", ".join(f"{self.env_name(n)}={self.knobs[n].render(values[n])}" for n in experiments)) + return KnobValues(self.prefix, values, sources) + + def serve_env(self, values: Optional[Mapping[str, Any]] = None) -> Dict[str, str]: + """``{"_": ""}`` for pinning in ``tt-model.yaml serve.env`` (defaults if no values).""" + vals = values if values is not None else self.defaults() + return {self.env_name(n): k.render(vals[n]) for n, k in self.knobs.items()} + + def doc_table(self) -> str: + """A Markdown table of the knobs (for SERVING.md / OPT_REPORT.md).""" + rows = ["| variable | default | meaning |", "|---|---|---|"] + for name, knob in self.knobs.items(): + extra = f" (one of {', '.join(map(str, knob.choices))})" if knob.choices else "" + tag = " **experiment**" if knob.experiment else "" + rows.append(f"| `{self.env_name(name)}` | `{knob.render(knob.default)}` | {knob.doc}{extra}{tag} |") + return "\n".join(rows) diff --git a/code/tt_diffusion_planner/ttaw/metrics.py b/code/tt_diffusion_planner/ttaw/metrics.py new file mode 100644 index 0000000000000000000000000000000000000000..c48539b2deb99b5631aa1e97baba586146bd8240 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/metrics.py @@ -0,0 +1,281 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C06 metrics: device-vs-reference agreement measures (host side, numpy only, float64 arithmetic). + +Inputs may be numpy arrays, torch tensors (bf16 included) or nested lists. Gates are applied by +:mod:`.golden` (``GateRegistry``); these functions only measure. + +========================== ===================================================================================== +function use +========================== ===================================================================================== +``pcc`` dense Pearson correlation (the default gate metric) +``masked_pcc`` PCC over the elements where a mask is true (valid pillars, in-range points) +``valid_row_pcc`` PCC over the first ``n_valid`` rows of a fixed-capacity buffer +``error_stats`` max / mean abs error, relative L2, PCC in one dict +``argmax_agreement`` fraction of positions whose argmax over ``axis`` agrees (segmentation logits) +``label_agreement`` fraction of equal labels (decided class maps) +``mask_iou`` ``mean_iou`` binary-mask IoU; per-class and mean IoU of label maps +``box_iou_xyxy`` pairwise IoU of 2-D boxes +``topk_overlap`` |A ∩ B| / k of two index sets (data-dependent selections) +``topk_set_overlap`` the same, computed from two score arrays +``match_detections`` greedy same-label matching by BEV-centre distance: recall, precision, errors +``ade_fde`` average / final displacement error of trajectories +========================== ===================================================================================== +""" +from __future__ import annotations + +import math +from dataclasses import dataclass, field +from typing import Any, Dict, List, Optional, Tuple + +import numpy as np + +__all__ = [ + "as_array", + "pcc", + "masked_pcc", + "valid_row_pcc", + "error_stats", + "argmax_agreement", + "label_agreement", + "mask_iou", + "mean_iou", + "box_iou_xyxy", + "topk_overlap", + "topk_set_overlap", + "DetectionMatch", + "match_detections", + "ade_fde", +] + + +def as_array(x: Any, dtype: Optional[Any] = None) -> np.ndarray: + """numpy / torch (any dtype, bf16 widened to float32) / list -> numpy, without importing torch.""" + if hasattr(x, "detach") and hasattr(x, "cpu"): + t = x.detach().cpu() + if str(t.dtype) == "torch.bfloat16": + t = t.float() + x = t.numpy() + a = np.asarray(x) + return a if dtype is None else a.astype(dtype, copy=False) + + +def _pcc_flat(a: np.ndarray, b: np.ndarray) -> float: + if a.shape != b.shape: + raise ValueError(f"shape mismatch: {a.shape} vs {b.shape}") + if a.size == 0: + return float("nan") + if not (np.isfinite(a).all() and np.isfinite(b).all()): + return float("nan") + # A constant side is detected exactly (max == min), never from the centred sum of squares: the mean of a + # constant array is not exact in floating point (np.full(1000, 0.1).mean() != 0.1), so the centred values hold + # rounding noise that would correlate perfectly with the other side's noise (two different constants gave +-1). + const_a, const_b = a.max() == a.min(), b.max() == b.min() + if const_a or const_b: # equal constants agree perfectly, anything else does not + return 1.0 if (const_a and const_b and np.allclose(a, b)) else 0.0 + da = a - a.mean() + db = b - b.mean() + sa = float(np.dot(da, da)) + sb = float(np.dot(db, db)) + return float(np.dot(da, db) / math.sqrt(sa * sb)) + + +def pcc(test: Any, ref: Any) -> float: + """Pearson correlation of two equally shaped tensors (float64). NaN if any value is non-finite or both are + empty; when a side is constant: 1.0 for two equal constants (``np.allclose``), else 0.0 (two different + constants included).""" + return _pcc_flat(as_array(test, np.float64).reshape(-1), as_array(ref, np.float64).reshape(-1)) + + +def masked_pcc(test: Any, ref: Any, mask: Any) -> float: + """PCC over ``mask`` (broadcastable to the tensors' shape).""" + a, b = as_array(test, np.float64), as_array(ref, np.float64) + if a.shape != b.shape: + raise ValueError(f"shape mismatch: {a.shape} vs {b.shape}") + m = np.broadcast_to(as_array(mask).astype(bool), a.shape) + return _pcc_flat(a[m], b[m]) + + +def valid_row_pcc(test: Any, ref: Any, n_valid: Optional[int] = None, *, rows: Any = None, axis: int = 0) -> float: + """PCC over the valid rows of fixed-capacity buffers: the first ``n_valid`` along ``axis``, or the rows where + the boolean vector ``rows`` is true.""" + a, b = as_array(test, np.float64), as_array(ref, np.float64) + if (n_valid is None) == (rows is None): + raise ValueError("give exactly one of n_valid and rows") + if rows is None: + idx = np.arange(int(n_valid)) + else: + idx = np.flatnonzero(as_array(rows).astype(bool)) + return _pcc_flat(np.take(a, idx, axis=axis).reshape(-1), np.take(b, idx, axis=axis).reshape(-1)) + + +def error_stats(test: Any, ref: Any) -> Dict[str, float]: + """``{"pcc", "max_abs", "mean_abs", "rel_l2", "n"}`` (``rel_l2 = ||t - r|| / ||r||``).""" + a, b = as_array(test, np.float64), as_array(ref, np.float64) + if a.shape != b.shape: + raise ValueError(f"shape mismatch: {a.shape} vs {b.shape}") + d = np.abs(a - b) + ref_norm = float(np.linalg.norm(b)) + return {"pcc": _pcc_flat(a.reshape(-1), b.reshape(-1)), + "max_abs": float(d.max()) if d.size else 0.0, + "mean_abs": float(d.mean()) if d.size else 0.0, + "rel_l2": float(np.linalg.norm(a - b) / ref_norm) if ref_norm > 0 else float(np.linalg.norm(a - b)), + "n": int(a.size)} + + +def argmax_agreement(test_logits: Any, ref_logits: Any, *, axis: int = -1, mask: Any = None) -> float: + """Fraction of positions where ``argmax`` over ``axis`` agrees (ties resolve to the lowest index, as numpy + and Autoware's decision rules do). ``mask`` selects positions (shape without ``axis``).""" + a = np.argmax(as_array(test_logits), axis=axis) + b = np.argmax(as_array(ref_logits), axis=axis) + return label_agreement(a, b, mask=mask) + + +def label_agreement(test_labels: Any, ref_labels: Any, *, mask: Any = None) -> float: + """Fraction of equal labels (optionally over ``mask``). NaN when nothing is selected.""" + a, b = as_array(test_labels), as_array(ref_labels) + if a.shape != b.shape: + raise ValueError(f"shape mismatch: {a.shape} vs {b.shape}") + eq = a == b + if mask is not None: + eq = eq[np.broadcast_to(as_array(mask).astype(bool), eq.shape)] + return float(eq.mean()) if eq.size else float("nan") + + +def mask_iou(test_mask: Any, ref_mask: Any) -> float: + """IoU of two binary masks (1.0 when both are empty).""" + a, b = as_array(test_mask).astype(bool), as_array(ref_mask).astype(bool) + union = np.logical_or(a, b).sum() + return 1.0 if union == 0 else float(np.logical_and(a, b).sum() / union) + + +def mean_iou(test_labels: Any, ref_labels: Any, num_classes: int, *, + ignore_index: Optional[int] = None) -> Tuple[float, List[float]]: + """``(mIoU over classes present in either map, per-class IoU with NaN for absent classes)``.""" + a, b = as_array(test_labels).reshape(-1), as_array(ref_labels).reshape(-1) + if a.shape != b.shape: + raise ValueError(f"shape mismatch: {a.shape} vs {b.shape}") + if ignore_index is not None: + keep = b != ignore_index + a, b = a[keep], b[keep] + per = [] + for c in range(num_classes): + inter = np.logical_and(a == c, b == c).sum() + union = np.logical_or(a == c, b == c).sum() + per.append(float(inter / union) if union else float("nan")) + valid = [v for v in per if not math.isnan(v)] + return (float(np.mean(valid)) if valid else float("nan")), per + + +def box_iou_xyxy(boxes_a: Any, boxes_b: Any) -> np.ndarray: + """Pairwise IoU ``[N, M]`` of axis-aligned boxes ``[x1, y1, x2, y2]`` (float64).""" + a = as_array(boxes_a, np.float64).reshape(-1, 4) + b = as_array(boxes_b, np.float64).reshape(-1, 4) + ix1 = np.maximum(a[:, None, 0], b[None, :, 0]) + iy1 = np.maximum(a[:, None, 1], b[None, :, 1]) + ix2 = np.minimum(a[:, None, 2], b[None, :, 2]) + iy2 = np.minimum(a[:, None, 3], b[None, :, 3]) + inter = np.clip(ix2 - ix1, 0, None) * np.clip(iy2 - iy1, 0, None) + area_a = np.clip(a[:, 2] - a[:, 0], 0, None) * np.clip(a[:, 3] - a[:, 1], 0, None) + area_b = np.clip(b[:, 2] - b[:, 0], 0, None) * np.clip(b[:, 3] - b[:, 1], 0, None) + union = area_a[:, None] + area_b[None, :] - inter + with np.errstate(invalid="ignore", divide="ignore"): + return np.where(union > 0, inter / union, 0.0) + + +def topk_overlap(test_indices: Any, ref_indices: Any, k: Optional[int] = None) -> float: + """``|set(test[:k]) ∩ set(ref[:k])| / k`` (``k`` defaults to ``len(ref)``).""" + a, b = as_array(test_indices).reshape(-1), as_array(ref_indices).reshape(-1) + k = len(b) if k is None else int(k) + if k <= 0: + return float("nan") + return len(set(a[:k].tolist()) & set(b[:k].tolist())) / k + + +def topk_set_overlap(test_scores: Any, ref_scores: Any, k: int) -> float: + """Overlap of the top-``k`` index sets of two score arrays (flattened; stable descending order).""" + a, b = as_array(test_scores, np.float64).reshape(-1), as_array(ref_scores, np.float64).reshape(-1) + ta = np.argsort(-a, kind="stable")[:k] + tb = np.argsort(-b, kind="stable")[:k] + return topk_overlap(ta, tb, k) + + +@dataclass +class DetectionMatch: + """Result of :func:`match_detections` (``pairs`` are ``(test_index, ref_index)``).""" + + pairs: List[Tuple[int, int]] + n_test: int + n_ref: int + center_errors: List[float] = field(default_factory=list) + score_errors: List[float] = field(default_factory=list) + + @property + def matched(self) -> int: + return len(self.pairs) + + @property + def recall(self) -> float: + return self.matched / self.n_ref if self.n_ref else 1.0 + + @property + def precision(self) -> float: + return self.matched / self.n_test if self.n_test else 1.0 + + def to_dict(self) -> Dict[str, Any]: + def stat(v: List[float], fn) -> Optional[float]: + return float(fn(v)) if v else None + + return {"n_test": self.n_test, "n_ref": self.n_ref, "matched": self.matched, "recall": self.recall, + "precision": self.precision, "center_err_mean": stat(self.center_errors, np.mean), + "center_err_max": stat(self.center_errors, np.max), "score_err_max": stat(self.score_errors, np.max), + "score_err_mean": stat(self.score_errors, np.mean)} + + +def _centers(x: Any, n: int, dims: int) -> np.ndarray: + a = as_array(x, np.float64) + if a.size == 0: + return np.zeros((n, dims)) + if a.size % max(n, 1) or n == 0: + raise ValueError(f"{a.shape[0] if a.ndim else 0} centres for {n} labels: lengths differ") + return a.reshape(n, -1)[:, :dims] + + +def match_detections(test_centers: Any, test_labels: Any, test_scores: Any, ref_centers: Any, ref_labels: Any, + ref_scores: Any, *, max_dist: float = 0.5, dims: int = 2, + same_label: bool = True) -> DetectionMatch: + """Greedy matching: reference boxes in descending score order each take the nearest unused test box (same + label unless ``same_label=False``) within ``max_dist`` of its centre, measured over the first ``dims`` + coordinates (2 = BEV x/y). ``*_centers`` are ``[N, >=dims]`` (box arrays ``[N, 7]`` work as-is). A box with a + non-finite centre never matches (it costs recall / precision; it does not block the other boxes).""" + tl, rl = as_array(test_labels).reshape(-1), as_array(ref_labels).reshape(-1) + ts, rs = as_array(test_scores, np.float64).reshape(-1), as_array(ref_scores, np.float64).reshape(-1) + if len(ts) != len(tl) or len(rs) != len(rl): + raise ValueError(f"labels / scores lengths differ: test {len(tl)}/{len(ts)}, ref {len(rl)}/{len(rs)}") + tc, rc = _centers(test_centers, len(tl), dims), _centers(ref_centers, len(rl), dims) + used = np.zeros(len(tl), bool) + result = DetectionMatch([], len(tl), len(rl)) + for i in np.argsort(-rs, kind="stable"): + if not len(tl): + break + d = np.linalg.norm(tc - rc[i], axis=1) + bad = used | ~np.isfinite(d) # np.argmin would pick a NaN distance before any real one + if same_label: + bad |= tl != rl[i] + d[bad] = np.inf + j = int(np.argmin(d)) + if d[j] <= max_dist: + used[j] = True + result.pairs.append((j, int(i))) + result.center_errors.append(float(d[j])) + result.score_errors.append(abs(float(ts[j] - rs[i]))) + return result + + +def ade_fde(pred: Any, ref: Any, *, dims: int = 2) -> Tuple[float, float]: + """Average and final displacement errors of trajectories ``[..., T, >=dims]`` (Euclidean over the first + ``dims`` coordinates; averaged over all leading dims).""" + p, r = as_array(pred, np.float64), as_array(ref, np.float64) + if p.shape[:-1] != r.shape[:-1]: + raise ValueError(f"trajectory shapes differ: {p.shape} vs {r.shape}") + disp = np.linalg.norm(p[..., :dims] - r[..., :dims], axis=-1) + return float(disp.mean()), float(disp[..., -1].mean()) diff --git a/code/tt_diffusion_planner/ttaw/models/__init__.py b/code/tt_diffusion_planner/ttaw/models/__init__.py new file mode 100644 index 0000000000000000000000000000000000000000..2fd611f43dbd1e038726ea1e660cf4c4bcbe36b8 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/models/__init__.py @@ -0,0 +1,24 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Shared device sub-networks of the Autoware ports (PLAN.md section 1.2), built from host (numpy) weights with +BatchNorm already folded, on the C17 builders (``ttaw.ops.conv``: ``FeatureMap``, ``Conv2d``, ``KSplitConv``, +``ConvTranspose2d``); every op traceable, weights prepared by the first (eager) call. + +================== ====== ========================================================================= +module id what it provides +================== ====== ========================================================================= +``second`` C24 SECOND backbone and SECONDFPN neck (CenterPoint, PointPainting, TransFusion, + BEVFusion); ``RowLinear`` (1x1 convs as ``ttnn.linear`` on NHWC rows) +``centerhead`` C25 CenterPoint dense head: K-split shared conv, merged sibling heads, one packed + fp32 output +``pillars`` C24 PillarFeatureNet of the pillar detectors (CenterPoint, PointPainting, TransFusion): + identity-folded split PFN on 32-slot pillar tiles, input packing, ``FrameStaging`` +``resnet`` C27 ResNet bottleneck / basic-block builders from a parameter provider, cameras as + batch (BEVDet image backbone + CustomResNet; BEVFormer, METEOR) +``transfusion_head`` C26 TransFusion query head (TransFusion, BEVFusion, PTv3): C21 selection, query init + with QPE / KPE tables, one decoder layer on C20, merged prediction heads +================== ====== ========================================================================= + +Importing a module never imports ttnn / torch (rule of the package): device work happens in the layer calls. +""" + +__all__ = ["second", "centerhead", "pillars", "resnet", "transfusion_head"] diff --git a/code/tt_diffusion_planner/ttaw/models/centerhead.py b/code/tt_diffusion_planner/ttaw/models/centerhead.py new file mode 100644 index 0000000000000000000000000000000000000000..4e25b71096b502e02230aba8bfa2b9bc1522821e --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/models/centerhead.py @@ -0,0 +1,164 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C25: the CenterPoint dense head on the device (PLAN.md section 1.2 row C25; CenterPoint, PointPainting). + +mmdet3d ``CenterHead`` (single task, as Autoware's exports): ``shared = relu(conv3x3(concat[d_0..d_n]))``, then per +output head ``out_h = conv1x1(relu(conv1x1(shared)))``. Two exact rewrites (each differs from the literal graph +only by summation order; probe P14 and research/centerpoint/SPEC.md 8.1): + +- **K-split shared conv** (default; C17 ``KSplitConv``): ``conv(concat[d_i], W) == sum_i conv(d_i, W_i)``, the bias + on the first partial conv and the ReLU fused into the last add. The concat (CenterPoint base: 177 MB) never + exists; P14: 3.6 ms against 14.8 ms for the full-K conv, and more accurate. ``accumulate_dtype="float32"`` keeps + the partials in fp32; +- **merged sibling heads**: the six 1x1 -> ReLU -> 1x1 heads become two matmuls on NHWC rows (:class:`..second. + RowLinear`): ``hidden = relu(x @ Wh + bh)`` with ``Wh = [Wh_1 | ... | Wh_6]`` (``64 -> 6*64``), then ``out = + hidden @ Wo + bo`` with the block-diagonal ``Wo`` (``6*64 -> sum C_h``, zero-padded to a multiple of 32 columns so + the output tensor is tile-exact). + +The output is ONE ``[1, 1, H*W, C_pad]`` tensor (fp32 by default: the logits feed host thresholds), i.e. one +readback; :meth:`CenterHead.unpack` splits the host rows into ``{name: (C, H, W)}``. Module names for the precision +policy: ``.shared``, ``.hidden``, ``.out``. +""" +from __future__ import annotations + +from typing import Any, Dict, List, Optional, Sequence, Tuple + +import numpy as np + +from ..ops.conv import Conv2d, FeatureMap, KSplitConv +from ..precision import PrecisionPolicy +from .second import ConvSpec, RowLinear, _precision, _to_dram, free_maps + +__all__ = ["CenterHead", "HeadSpec", "merge_heads", "centerhead_numpy"] + +HeadSpec = Tuple[str, ConvSpec, ConvSpec] # (output name, hidden 1x1 + ReLU, final 1x1) + + +def merge_heads(heads: Sequence[HeadSpec], *, pad_to: int = 32 + ) -> Tuple[np.ndarray, np.ndarray, np.ndarray, np.ndarray, List[Tuple[str, int, int]]]: + """``(Wh [Cin, sum hidden], bh, Wo [sum hidden, C_pad], bo [C_pad], slices)`` of the merged heads; ``slices`` + lists ``(name, start, stop)`` of each head's output columns; columns past ``sum C_h`` are zero.""" + if not heads: + raise ValueError("no heads") + cin = heads[0][1].in_channels + wh, bh, rows, cols = [], [], 0, 0 + for name, hidden, final in heads: + if hidden.kind != "conv" or final.kind != "conv" or hidden.kernel != 1 or final.kernel != 1: + raise ValueError(f"head {name!r}: merged heads need 1x1 convs") + if hidden.in_channels != cin or final.in_channels != hidden.out_channels: + raise ValueError(f"head {name!r}: channel mismatch") + if not hidden.relu or final.relu: + raise ValueError(f"head {name!r}: expected a ReLU after the hidden conv only") + wh.append(np.asarray(hidden.weight, np.float32)[:, :, 0, 0].T) + bh.append(np.zeros(hidden.out_channels, np.float32) if hidden.bias is None + else np.asarray(hidden.bias, np.float32)) + rows += hidden.out_channels + cols += final.out_channels + cpad = -(-cols // pad_to) * pad_to + wo = np.zeros((rows, cpad), np.float32) + bo = np.zeros(cpad, np.float32) + slices, r, c = [], 0, 0 + for name, hidden, final in heads: + co = final.out_channels + wo[r:r + hidden.out_channels, c:c + co] = np.asarray(final.weight, np.float32)[:, :, 0, 0].T + if final.bias is not None: + bo[c:c + co] = np.asarray(final.bias, np.float32) + slices.append((name, c, c + co)) + r += hidden.out_channels + c += co + return (np.ascontiguousarray(np.concatenate(wh, axis=1)), np.concatenate(bh).astype(np.float32), wo, bo, + slices) + + +class CenterHead: + """Shared conv (K-split over ``split``) + merged heads; see the module docstring. + + ``shared``: the 3x3 :class:`ConvSpec` over the concatenated inputs (ReLU after); ``split``: channels of each + input, in concat order; ``heads``: ``(name, hidden, final)`` in output order. ``k_split=False`` runs one conv + on a device concat instead (A/B only). ``output_dtype``: of the packed head output (``"float32"``). The shared + conv's output dtype is its precision's ``activations`` (bf16 by default); ``accumulate_dtype`` (default: the same) + is the dtype of the K-split partial sums.""" + + def __init__(self, shared: ConvSpec, split: Sequence[int], heads: Sequence[HeadSpec], *, + policy: Optional[PrecisionPolicy] = None, prefix: str = "head", output_dtype: str = "float32", + k_split: bool = True, accumulate_dtype: Optional[str] = None): + if shared.kind != "conv" or shared.stride != 1 or not shared.relu: + raise ValueError("the shared conv must be a stride-1 conv followed by a ReLU") + self.prefix = prefix + self.split = [int(s) for s in split] + if sum(self.split) != shared.in_channels: + raise ValueError(f"split {self.split} does not cover the {shared.in_channels} shared-conv inputs") + self.k_split = bool(k_split) and len(self.split) > 1 + pad = shared.kernel // 2 if shared.padding is None else int(shared.padding) + prec = _precision(policy, f"{prefix}.shared") + if self.k_split: + self.shared = KSplitConv(shared.weight, shared.bias, self.split, activation="relu", padding=pad, + precision=prec, accumulate_dtype=accumulate_dtype or prec.activations, + output_dtype=prec.activations, name=f"{prefix}.shared") + else: + self.shared = Conv2d(shared.weight, shared.bias, padding=pad, activation="relu", precision=prec, + output_dtype=prec.activations, name=f"{prefix}.shared") + wh, bh, wo, bo, self.slices = merge_heads(heads) + self.out_channels = self.slices[-1][2] + self.padded_channels = int(wo.shape[1]) + hidden_prec = _precision(policy, f"{prefix}.hidden") + self.hidden = RowLinear(wh, bh, activation="relu", precision=hidden_prec, + output_dtype=hidden_prec.activations, name=f"{prefix}.hidden") + self.out = RowLinear(wo, bo, activation=None, precision=_precision(policy, f"{prefix}.out"), + output_dtype=output_dtype, name=f"{prefix}.out") + + def shared_forward(self, inputs: Sequence[FeatureMap]) -> FeatureMap: + """``relu(conv3x3(concat(inputs)))`` (inputs are kept), DRAM interleaved.""" + import ttnn + + if [x.channels for x in inputs] != self.split: + raise ValueError(f"{self.prefix}: inputs have {[x.channels for x in inputs]} channels, " + f"expected {self.split}") + if self.k_split: + return _to_dram(self.shared(inputs)) + cat = inputs[0].with_tensor(ttnn.concat([x.tensor for x in inputs], dim=-1), sum(self.split)) + out = _to_dram(self.shared(cat)) + free_maps(cat) + return out + + def heads_forward(self, shared: FeatureMap) -> FeatureMap: + """Merged heads on the shared map -> ``[1, 1, H*W, C_pad]`` (``channels`` = the padded width).""" + hidden = self.hidden(shared.tensor) + out = self.out(hidden) + free_maps(hidden) + return shared.with_tensor(out, self.padded_channels) + + def __call__(self, inputs: Sequence[FeatureMap]) -> FeatureMap: + shared = self.shared_forward(inputs) + out = self.heads_forward(shared) + free_maps(shared) + return out + + def unpack(self, rows: Any, height: int, width: int) -> Dict[str, np.ndarray]: + """Host rows ``[H*W, >= C]`` (a read-back of the head output) -> ``{name: (C_h, H, W)}`` float32.""" + a = np.asarray(rows, dtype=np.float32) + a = a.reshape(-1, a.shape[-1])[: height * width] + return {name: np.ascontiguousarray(a[:, s:e].T.reshape(e - s, height, width)) for name, s, e in self.slices} + + def release(self) -> None: + self.shared.release() + self.hidden.release() + self.out.release() + + def describe(self) -> Dict[str, Any]: + return {"shared": self.shared.describe(), "k_split": self.k_split, "hidden": self.hidden.describe(), + "out": self.out.describe(), "slices": [list(s) for s in self.slices]} + + +def centerhead_numpy(shared: ConvSpec, heads: Sequence[HeadSpec], inputs_nchw: Sequence[Any]) -> Dict[str, np.ndarray]: + """fp32 oracle (torch, the literal graph: concat, conv, six heads) -> ``{"shared": NCHW, name: NCHW}``.""" + import torch + + from .second import _apply_torch + + with torch.no_grad(): + x = torch.cat([torch.as_tensor(np.asarray(t, dtype=np.float32)) for t in inputs_nchw], dim=1) + s = _apply_torch(shared, x) + out = {"shared": s.numpy().copy()} + for name, hidden, final in heads: + out[name] = _apply_torch(final, _apply_torch(hidden, s)).numpy().copy() + return out diff --git a/code/tt_diffusion_planner/ttaw/models/pillars.py b/code/tt_diffusion_planner/ttaw/models/pillars.py new file mode 100644 index 0000000000000000000000000000000000000000..9be24d7a56b7b637f548b1700ea899de73ae656e --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/models/pillars.py @@ -0,0 +1,297 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C24 companion: the PillarFeatureNet (PFN) of the pillar detectors on the device, and the host staging of its +inputs (PLAN.md section 1.2; CenterPoint, PointPainting, TransFusion). Promoted by the PointPainting port from the +CenterPoint bundle's ``tt/pfn.py`` + ``tt/model.py FrameStaging`` (device-verified there: PFN PCC 0.99997) with the +TransFusion bundle's generalisation (any hidden / output width, pillars of fewer than 32 points); those bundles keep +their copies until they re-vendor. + +The encoder (mmdet3d ``PFNLayer`` x 2 as Autoware exports it; BN folded into both MatMuls by the port's loader): + + h = relu(x @ W0 + b0) x: (P, K, F) decorated points, padded slots all zero h: (P, K, H) + m = max over the K slots of h padded slots included: their rows are relu(b0) + y = max over the slots of relu([h, m] @ W1 + b1) = relu(max_slots(h @ W1a) + m @ W1b + b1) y: (P, O) + +Device form. One pillar is one 32-row tile (its point slots), so a slot max is a tile-column reduction (``ttnn.max`` +over dim -2). Because ``relu(. + c)`` is monotonic, the per-pillar term ``c = m @ W1b + b1`` is added AFTER the max +(exact, also in floating point: rounding is monotonic), so no slot broadcast is needed; two identity blocks fold the +rest into the matmuls (exact: a value times 1.0 with HiFi4 and fp32 accumulation): + + x [P*32, 16] --linear W0 (16 x H_pad) + b0, ReLU--> h [P*32, H_pad] + h --linear Wz = [I_H | W1a] (H_pad x Z)--> z [P*32, Z] = [h | h @ W1a | 0] + z as [P, 32, Z] --max over dim -2--> mz [P, 1, Z] = [m | max(h @ W1a) | 0] + mz --linear Wy = [W1b ; I_O] (Z x O) + b1, ReLU--> y [P, 1, O] + y --untilize, view--> [1, 1, P, O] ROW_MAJOR bf16 rows + +(``H_pad`` = H rounded up to 32, ``Z`` = H + O rounded up to 32, ``O % 32 == 0``: the rows are a C19 gather table.) +The rows feed the gather-form scatter (C19 ``ops.gather.scatter_rows``, which appends the zero sentinel row). + +**Pillars of K < 32 points** (TransFusion: 20): slots K..31 repeat slot 0 (:func:`pack_features`; every kept +pillar has a point in slot 0), so both slot maxima are unchanged (max is idempotent), whereas zero rows there would +add ``relu(b0)`` to the maxima of full pillars. With K = 32 (CenterPoint, PointPainting) the padded slots are the +model's own zero slots. + +**Input**: the host's decorated features as bf16 ROW_MAJOR ``[1, 1, P, 32 * 16]``: one 1 KB row per pillar (32 +slots x F <= 16 features zero-padded to 16). One row per pillar keeps the host-to-device copy at P pages (CenterPoint: +5.1 ms for 40,000 pages of 1 KB vs 32 ms as 1.28 M pages of 32 B); the device reshapes it to ``[1, 1, P*32, 16]`` and +tilizes it. Every pillar row is computed, also past the frame's pillar count (stale or zero input): the scatter index +never points at those rows, so no pillar count reaches the device. :class:`FrameStaging` writes these inputs in place +into persistent host tensors (``tensors.HostStaging``): no per-frame ``from_torch`` of 20 M elements (CenterPoint: 188 +ms) and no allocation; only the kept pillars are converted (torch fp32 -> bf16, RNE). + +**Precision** (``ttaw.precision.Precision``): HiFi4 is required (the identity blocks are exact only with it); fp32 +accumulation; weights per ``Precision.weights`` (fp32: a device "fp32" matmul behaves like TF32, probe P12); +``Precision.activations`` is the dtype of the hidden tensors ``h`` and ``z`` (bf16 by default; fp32 keeps the PFN's +internal roundings off: PointPainting measured on the CPU that the bf16 rounding of ``h`` / ``z`` is the largest PFN +error term, rel. L2 0.0045 vs 0.0017 for the bf16 output rounding alone). The output rows are always bf16 (the +``ttnn.embedding`` table of C19 is bf16-only). + +numpy only at import; ttnn / torch are imported inside the device and staging functions. +""" +from __future__ import annotations + +from dataclasses import dataclass +from typing import Any, Dict, Optional + +import numpy as np + +from ..precision import Precision +from .second import RowLinear, free_maps + +__all__ = ["PFN_IN_FEATURES", "PILLAR_SLOTS", "PfnPlan", "TtPillarFeatureNet", "FrameStaging", "features_shape", + "pack_features", "pfn_reference_numpy", "pack_inputs", "DEFAULT_PRECISION"] + +PFN_IN_FEATURES = 16 # decorated features zero-padded to 16 bf16 = 32-byte slot rows +PILLAR_SLOTS = 32 # one 32-row tile per pillar +TILE = 32 +DEFAULT_PRECISION = "HiFi4+fp32:w=fp32" + + +def _round_up(n: int, m: int = TILE) -> int: + return -(-int(n) // m) * m + + +@dataclass(frozen=True) +class PfnPlan: + """The device matrices of the PFN (float32 host arrays; see the module docstring).""" + + w0: np.ndarray # (16, H_pad) + b0: np.ndarray # (H_pad,) + wz: np.ndarray # (H_pad, Z) = [I_H | W1a | 0] + wy: np.ndarray # (Z, O) = [W1b ; I_O ; 0] + b1: np.ndarray # (O,) + in_features: int # F + hidden: int # H (layer-0 width) + out_features: int # O + + @classmethod + def from_weights(cls, w0: Any, b0: Any, w1: Any, b1: Any, *, in_padded: int = PFN_IN_FEATURES) -> "PfnPlan": + """From the BN-folded encoder weights: ``w0`` (F, H) and ``w1`` (2H, O), ``[in, out]`` layout; layer 1 reads + ``concat[h, max_slots h]`` (rows 0..H-1 of ``w1`` act on ``h``, rows H..2H-1 on the max).""" + w0 = np.asarray(w0, np.float32) + w1 = np.asarray(w1, np.float32) + if w0.ndim != 2 or w1.ndim != 2: + raise ValueError(f"w0 {w0.shape} / w1 {w1.shape} must be [in, out] matrices") + f, hdim = w0.shape + out = int(w1.shape[1]) + if w1.shape[0] != 2 * hdim: + raise ValueError(f"w1 {w1.shape} does not take concat[h, max h] of width {2 * hdim}") + if f > in_padded: + raise ValueError(f"{f} features do not fit the {in_padded}-wide slot rows") + if out % TILE: + raise ValueError(f"PFN output width {out}: the C19 gather table needs a multiple of {TILE}") + b0 = np.asarray(b0, np.float32).reshape(-1) + b1 = np.asarray(b1, np.float32).reshape(-1) + if b0.shape != (hdim,) or b1.shape != (out,): + raise ValueError(f"biases {b0.shape} / {b1.shape} do not match widths {hdim} / {out}") + hp, zp = _round_up(hdim), _round_up(hdim + out) + W0 = np.zeros((in_padded, hp), np.float32) + W0[:f, :hdim] = w0 + B0 = np.zeros(hp, np.float32) + B0[:hdim] = b0 + Wz = np.zeros((hp, zp), np.float32) + Wz[:hdim, :hdim] = np.eye(hdim, dtype=np.float32) + Wz[:hdim, hdim:hdim + out] = w1[:hdim] + Wy = np.zeros((zp, out), np.float32) + Wy[:hdim] = w1[hdim:] + Wy[hdim:hdim + out] = np.eye(out, dtype=np.float32) + return cls(W0, B0, Wz, Wy, b1.copy(), int(f), int(hdim), out) + + def forward_numpy(self, features: Any, *, hidden_round: Any = None) -> np.ndarray: + """float32 host emulation of the device graph (``(P, S, F)`` slots -> ``(P, out_features)``); give it the + 32-slot layout of :func:`pack_features` to emulate the device's slots exactly. ``hidden_round``: a function + applied to ``h`` and ``z`` (e.g. a bf16 rounding) to model the hidden dtype.""" + r = hidden_round or (lambda a: a) + x = np.asarray(features, np.float32) + p, slots, f = x.shape + xp = np.zeros((p, slots, self.w0.shape[0]), np.float32) + xp[:, :, :f] = x + h = r(np.maximum(xp @ self.w0 + self.b0, 0.0).astype(np.float32)) + z = r((h @ self.wz).astype(np.float32)) + mz = z.max(axis=1) + return np.maximum(mz @ self.wy + self.b1, 0.0).astype(np.float32) + + +def pfn_reference_numpy(features: Any, w0: Any, b0: Any, w1: Any, b1: Any) -> np.ndarray: + """The literal encoder in float64 (concat form, max over the given slots, padded slots included) -> ``(P, O)`` + float32: the oracle of :class:`PfnPlan` and of the device module.""" + x = np.asarray(features, np.float64) + w0, w1 = np.asarray(w0, np.float64), np.asarray(w1, np.float64) + h = np.maximum(x @ w0 + np.asarray(b0, np.float64), 0.0) + m = h.max(axis=1, keepdims=True) + y = np.maximum(np.concatenate([h, np.broadcast_to(m, h.shape)], axis=2) @ w1 + np.asarray(b1, np.float64), 0.0) + return y.max(axis=1).astype(np.float32) + + +def features_shape(num_pillars: int): + """The device input layout of the PFN: ``(1, 1, num_pillars, 32 * 16)`` (one bf16 row per pillar).""" + return (1, 1, int(num_pillars), PILLAR_SLOTS * PFN_IN_FEATURES) + + +def pack_features(features: Any, num_pillars: int, *, out: Optional[np.ndarray] = None) -> np.ndarray: + """Decorated features ``(P_kept, K, F)`` (K <= 32 point slots, F <= 16; the model's padded slots zero) -> the + device input :func:`features_shape` as float32: features zero-padded to 16, slots K..31 copies of slot 0 when + K < 32 (module docstring), rows past ``P_kept`` zero. ``out`` (an array of that shape) is reused when given.""" + x = np.asarray(features, np.float32) + if x.ndim != 3: + raise ValueError(f"features must be (pillars, slots, features), got {x.shape}") + p, slots, f = x.shape + if slots > PILLAR_SLOTS or f > PFN_IN_FEATURES or p > num_pillars: + raise ValueError(f"features {x.shape} do not fit ({num_pillars}, {PILLAR_SLOTS}, <= {PFN_IN_FEATURES})") + shape = features_shape(num_pillars) + if out is None: + out = np.zeros(shape, np.float32) + elif out.shape != shape or out.dtype != np.float32: + raise ValueError(f"out must be float32 {shape}") + else: + out.fill(0.0) + v = out.reshape(int(num_pillars), PILLAR_SLOTS, PFN_IN_FEATURES) + v[:p, :slots, :f] = x + if 0 < slots < PILLAR_SLOTS: + v[:p, slots:, :f] = x[:, :1, :] + return out + + +def pack_inputs(features: Any, canvas_index: Any, num_pillars: int, *, capacity: int, + features_out: Optional[np.ndarray] = None) -> Dict[str, np.ndarray]: + """The two PFN-graph inputs as numpy (tests and tools; the API uses :class:`FrameStaging`): ``{"features": + (1, 1, capacity, 512) float32, "canvas_index": (1, N) uint32}`` from the kept pillars' features + ``features[:num_pillars]`` and a cell -> pillar-row index (``capacity`` = the zero sentinel row). The index is + range-checked: a bad index would read outside the gather table.""" + from ..ops.gather import check_index + + n = int(num_pillars) + feats = pack_features(np.asarray(features)[:n], capacity, out=features_out) + return {"features": feats, "canvas_index": check_index(canvas_index, num_rows=n, sentinel=capacity)} + + +class TtPillarFeatureNet: + """The PFN on ``num_pillars`` pillars: ``[1, 1, P, 32*16]`` bf16 ROW_MAJOR -> ``[1, 1, P, O]`` bf16 ROW_MAJOR + rows (module docstring). ``precision``: a :class:`..precision.Precision` or its spec; its ``activations`` dtype is + that of the hidden tensors ``h`` / ``z``.""" + + def __init__(self, plan: PfnPlan, *, num_pillars: int, precision: Any = DEFAULT_PRECISION, name: str = "pfn"): + self.plan = plan + self.name = name + self.num_pillars = int(num_pillars) + self.precision = Precision.parse(precision) + if self.precision.fidelity != "HiFi4": + raise ValueError(f"{name}: the PFN identity blocks are exact only with HiFi4 (x * 1.0)") + hidden = self.precision.activations + self.l0 = RowLinear(plan.w0, plan.b0, activation="relu", precision=self.precision, output_dtype=hidden, + name=f"{name}.l0") + self.lz = RowLinear(plan.wz, None, precision=self.precision, output_dtype=hidden, name=f"{name}.lz") + self.ly = RowLinear(plan.wy, plan.b1, activation="relu", precision=self.precision, output_dtype="bfloat16", + name=f"{name}.ly") + self.out_features = plan.out_features + + @property + def features_shape(self): + return features_shape(self.num_pillars) + + def __call__(self, features_rm: Any) -> Any: + import ttnn + + p = self.num_pillars + rows = ttnn.reshape(features_rm, (1, 1, p * PILLAR_SLOTS, PFN_IN_FEATURES)) # 1 KB rows -> 32 B slot rows + x = ttnn.to_layout(rows, ttnn.TILE_LAYOUT) # [1, 1, P*32, 16] + free_maps(rows) + h = self.l0(x) # [1, 1, P*32, H_pad] + free_maps(x) + z = self.lz(h) # [1, 1, P*32, Z] + free_maps(h) + z4 = ttnn.reshape(z, (1, p, PILLAR_SLOTS, int(z.shape[-1]))) # one pillar = one tile row block (a view) + mz = ttnn.max(z4, dim=2, keepdim=True) # [1, P, 1, Z] + free_maps(z) + y = self.ly(mz) # [1, P, 1, O] bf16 + free_maps(mz) + y_rm = ttnn.to_layout(y, ttnn.ROW_MAJOR_LAYOUT) + free_maps(y) + return ttnn.reshape(y_rm, (1, 1, p, self.out_features)) + + def release(self) -> None: + for layer in (self.l0, self.lz, self.ly): + layer.release() + + def describe(self) -> Dict[str, Any]: + return {"name": self.name, "num_pillars": self.num_pillars, "in_features": self.plan.in_features, + "hidden": self.plan.hidden, "out": self.plan.out_features, "z_width": int(self.plan.wz.shape[1]), + "slots": PILLAR_SLOTS, "precision": self.precision.label, + "hidden_dtype": self.precision.activations} + + +class FrameStaging: + """Persistent host buffers of the two PFN-graph inputs: bf16 ``features`` (:func:`features_shape`) and uint32 + ``canvas_index`` ``[1, num_cells]`` (``tensors.HostStaging``: the ttnn host tensor's own memory, written in + place). :meth:`write` converts only the kept pillars (torch fp32 -> bf16, RNE) into their rows (slots K..31 = + slot 0 when K < 32), zeroes the rows the previous frame used beyond them, copies the range-checked index, and + returns the ``TraceRunner`` inputs (host tensors, uploaded as they are). Without a zero-copy alias (another + tt-metal build) it falls back to :func:`pack_features` + ``HostStaging.write`` (same values).""" + + def __init__(self, num_pillars: int, num_cells: int): + from ..tensors import HostStaging + + self.num_pillars = int(num_pillars) + self.num_cells = int(num_cells) + self.features = HostStaging(features_shape(self.num_pillars), "bfloat16") + self.index = HostStaging((1, self.num_cells), "uint32") + self._used = self.num_pillars # rows that may hold data (all, until the first write) + + @property + def zero_copy(self) -> bool: + return bool(self.features.zero_copy and self.index.zero_copy) + + def write(self, features: Any, canvas_index: Any, num_pillars: Optional[int] = None) -> Dict[str, Any]: + """``features``: ``(P, K, F)`` float32 decorated pillars (only the first ``num_pillars`` rows are used; + default all); ``canvas_index``: ``num_cells`` entries in ``[0, num_pillars)`` or the sentinel + ``self.num_pillars``.""" + from ..ops.gather import check_index + + f_all = np.asarray(features) + p = int(f_all.shape[0] if num_pillars is None else num_pillars) + if p > self.num_pillars: + raise ValueError(f"{p} pillars exceed the staging capacity {self.num_pillars}") + feats = np.asarray(f_all[:p], np.float32) + index = check_index(canvas_index, num_rows=p, sentinel=self.num_pillars) + if index.size != self.num_cells: + raise ValueError(f"canvas index of {index.size} cells, expected {self.num_cells}") + if not self.zero_copy: # no host alias: plain writes + return {"features": self.features.write(pack_features(feats, self.num_pillars)), + "canvas_index": self.index.write(index)} + import torch + + if feats.ndim != 3 or feats.shape[1] > PILLAR_SLOTS or feats.shape[2] > PFN_IN_FEATURES: + raise ValueError(f"features {feats.shape} do not fit (P, <= {PILLAR_SLOTS}, <= {PFN_IN_FEATURES})") + view = self.features.view.reshape(self.num_pillars, PILLAR_SLOTS, PFN_IN_FEATURES).view(np.int16) + dst = torch.from_numpy(view) + slots, f = feats.shape[1], feats.shape[2] + src = torch.from_numpy(np.ascontiguousarray(feats)).to(torch.bfloat16).view(torch.int16) + dst[:p, :slots, :f] = src + if f < PFN_IN_FEATURES: + dst[:p, :, f:] = 0 + if 0 < slots < PILLAR_SLOTS: + dst[:p, slots:, :f] = src[:, :1, :] + if p < self._used: + dst[p:self._used] = 0 + self._used = p + self.index.view[...] = index + return {"features": self.features.tensor, "canvas_index": self.index.tensor} diff --git a/code/tt_diffusion_planner/ttaw/models/resnet.py b/code/tt_diffusion_planner/ttaw/models/resnet.py new file mode 100644 index 0000000000000000000000000000000000000000..c264e6bc87ad4ef00c0686b9bc925d5d283435a4 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/models/resnet.py @@ -0,0 +1,271 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C27: ResNet builders on device feature maps, cameras as batch (PLAN.md section 1.2 row C27; first built by the +BEVDet port for its ResNet-50 image backbone and CustomResNet BEV encoder; BEVFormer and METEOR reuse it). + +- :class:`Bottleneck`: ResNet-50/101, pytorch style (the stride sits on the 3x3 ``conv2``; 1x1 projection + ``downsample`` on the first block of a stage): ``relu(conv3(relu(conv2(relu(conv1(x))))) + idt)``; +- :class:`BasicBlock`: ResNet-18/34 and BEVDet's CustomResNet (two 3x3 convs; the projection may be any conv, + CustomResNet's is a 3x3 stride-2 one): ``relu(conv2(relu(conv1(x))) + idt)``; +- :class:`ResNetStem`: ``maxpool 3x3 s2 p1 (relu(conv1 7x7 s2 p3))`` (the max pool runs DRAM-sliced); +- :class:`ResNetStages`: ``len(layers)`` stages of one block kind; ``__call__(x, keep=...)`` returns the stage + outputs a caller wants (FPN levels, taps) and frees the others as soon as the next stage consumed them. + +**Weights** come from a parameter provider ``provider(module) -> ConvParams`` (BatchNorm folded, conv attributes as +the source graph defines them: BEVDet's reads ``OnnxWeights`` by consuming node, so no kernel size or stride is +re-typed here). **Names** come from a ``names(stage, block)`` callable: :func:`mmdet_names` (``layer1.0.conv1``, +``downsample.0``) and :func:`custom_resnet_names` (``layers.0.0.conv1``, ``downsample``). **Convs** come from a +``factory(module, params, activation)`` (default :func:`conv_factory`: one C17 ``Conv2d`` per module, with a +``ttaw.precision`` spec or a ``PrecisionPolicy`` resolved per module name); a port that needs deformable convs +(BEVFormer / METEOR DCN, PLAN.md C23 / C27 "DCN hook") returns its own callable for the modules it names. + +Placement: every activation is a DRAM-interleaved TILE ``[1, 1, N*H*W, C]`` map (C17's rule: DRAM inputs run ttnn's +automatic DRAM slicing, 1x1 convs are matmuls placed back in DRAM), so the 6 x 256 x 704 cameras' 34.6 MB stem and +layer1 maps fit; intermediates are freed as soon as they are consumed. Every op is traceable; a layer prepares its +weights on its first (eager) call. :func:`resnet_numpy` is the fp32 torch oracle of the same blocks. +""" +from __future__ import annotations + +from dataclasses import dataclass +from typing import Any, Callable, Dict, List, Optional, Sequence, Tuple, Union + +import numpy as np + +from ..ops.conv import CONV_PRECISION, Conv2d, FeatureMap, residual_add +from ..precision import Precision, PrecisionPolicy + +__all__ = ["ConvParams", "DEFAULT_PRECISION", "conv_factory", "mmdet_names", "custom_resnet_names", "Bottleneck", + "BasicBlock", "ResNetStem", "ResNetStages", "maxpool_out_hw", "resnet_numpy", "identity_like"] + +DEFAULT_PRECISION = CONV_PRECISION + + +@dataclass(frozen=True) +class ConvParams: + """One BN-folded convolution as a provider returns it (numpy float32).""" + + weight: np.ndarray # [Cout, Cin / groups, kh, kw] + bias: Optional[np.ndarray] = None # [Cout] + stride: Tuple[int, int] = (1, 1) + padding: Tuple[int, int, int, int] = (0, 0, 0, 0) # top, bottom, left, right + dilation: Tuple[int, int] = (1, 1) + groups: int = 1 + + +Provider = Callable[[str], ConvParams] +ConvFactory = Callable[[str, ConvParams, Optional[str]], Callable[[FeatureMap], FeatureMap]] +Names = Callable[[int, int], Dict[str, str]] + + +def conv_factory(precision: Union[str, Precision, PrecisionPolicy] = DEFAULT_PRECISION, **conv_kwargs: Any + ) -> ConvFactory: + """The standard factory: one C17 ``Conv2d`` per module. ``precision``: a spec / :class:`Precision` for every + conv, or a :class:`PrecisionPolicy` resolved per module name; other keywords go to every ``Conv2d``.""" + + def make(module: str, p: ConvParams, activation: Optional[str]) -> Conv2d: + prec = precision.resolve(module) if isinstance(precision, PrecisionPolicy) else precision + return Conv2d(p.weight, p.bias, stride=p.stride, padding=p.padding, dilation=p.dilation, groups=p.groups, + activation=activation, precision=prec, name=module, **conv_kwargs) + + return make + + +def mmdet_names(prefix: str) -> Names: + """mmdet / torchvision names: ``{prefix}.layer{stage + 1}.{block}.conv{1,2,3}`` and ``.downsample.0``.""" + + def names(stage: int, block: int) -> Dict[str, str]: + p = f"{prefix}.layer{stage + 1}.{block}" + return {"conv1": p + ".conv1", "conv2": p + ".conv2", "conv3": p + ".conv3", + "downsample": p + ".downsample.0", "block": p} + + return names + + +def custom_resnet_names(prefix: str) -> Names: + """BEVDet CustomResNet names: ``{prefix}.layers.{stage}.{block}.conv{1,2}`` and ``.downsample``.""" + + def names(stage: int, block: int) -> Dict[str, str]: + p = f"{prefix}.layers.{stage}.{block}" + return {"conv1": p + ".conv1", "conv2": p + ".conv2", "downsample": p + ".downsample", "block": p} + + return names + + +def identity_like(x: FeatureMap, like: FeatureMap) -> FeatureMap: + """The identity shortcut ``x`` in the tensor shape of the residual branch output ``like`` (same rows and + channels). ``ttnn.max_pool2d`` on a DRAM input (:class:`ResNetStem`) returns a 4-D ``[N, OH, OW, C]`` TILE map + (``generic_pools.cpp`` ``pool2d_DRAM``) while a conv returns the flattened ``[1, 1, N*OH*OW, C]`` rows, and + ``ttnn.add`` of the two raises ``Invalid subtile broadcast type``: the first block of a ResNet-18 / 34 stage 1 (no + projection) adds the stem output itself (METEOR port, ttaw 0.17.2). A no-op when the shapes already agree.""" + if tuple(x.tensor.shape) == tuple(like.tensor.shape): + return x + import ttnn + + return x.with_tensor(ttnn.reshape(x.tensor, tuple(like.tensor.shape))) + + +class Bottleneck: + """``relu(conv3(relu(conv2(relu(conv1(x))))) + idt)``; ``idt = downsample(x)`` when the block has a projection.""" + + def __init__(self, provider: Provider, names: Dict[str, str], *, has_downsample: bool, factory: ConvFactory): + self.name = names["block"] + self.conv1 = factory(names["conv1"], provider(names["conv1"]), "relu") + self.conv2 = factory(names["conv2"], provider(names["conv2"]), "relu") + self.conv3 = factory(names["conv3"], provider(names["conv3"]), None) + self.downsample = (factory(names["downsample"], provider(names["downsample"]), None) + if has_downsample else None) + + def __call__(self, x: FeatureMap) -> FeatureMap: + y = self.conv1(x) + z = self.conv2(y) + y.deallocate() + y = self.conv3(z) + z.deallocate() + if self.downsample is not None: + idt = self.downsample(x) + out = residual_add(y, idt, "relu") + idt.deallocate() + else: + out = residual_add(y, identity_like(x, y), "relu") + y.deallocate() + return out + + def convs(self) -> List[Any]: + return [c for c in (self.conv1, self.conv2, self.conv3, self.downsample) if c is not None] + + +class BasicBlock: + """``relu(conv2(relu(conv1(x))) + idt)``; ``idt = downsample(x)`` when the block has a projection.""" + + def __init__(self, provider: Provider, names: Dict[str, str], *, has_downsample: bool, factory: ConvFactory): + self.name = names["block"] + self.conv1 = factory(names["conv1"], provider(names["conv1"]), "relu") + self.conv2 = factory(names["conv2"], provider(names["conv2"]), None) + self.downsample = (factory(names["downsample"], provider(names["downsample"]), None) + if has_downsample else None) + + def __call__(self, x: FeatureMap) -> FeatureMap: + y = self.conv1(x) + z = self.conv2(y) + y.deallocate() + if self.downsample is not None: + idt = self.downsample(x) + out = residual_add(z, idt, "relu") + idt.deallocate() + else: + out = residual_add(z, identity_like(x, z), "relu") + z.deallocate() + return out + + def convs(self) -> List[Any]: + return [c for c in (self.conv1, self.conv2, self.downsample) if c is not None] + + +def maxpool_out_hw(h: int, w: int, k: int = 3, s: int = 2, p: int = 1) -> Tuple[int, int]: + return (h + 2 * p - k) // s + 1, (w + 2 * p - k) // s + 1 + + +class ResNetStem: + """``maxpool 3x3 s2 p1 (relu(conv1))`` (``conv_module`` names the stem conv; ``pool=False`` skips the pool). + The max pool reads the DRAM conv output with automatic DRAM slicing and writes a TILE map for the first block; + on that DRAM path its tensor is 4-D ``[N, OH, OW, C]`` (not the flattened rows of a conv output): convs take it + as is, and the blocks' identity shortcuts reshape it (:func:`identity_like`).""" + + def __init__(self, provider: Provider, conv_module: str, *, factory: Optional[ConvFactory] = None, + pool: bool = True): + factory = factory or conv_factory() + self.conv1 = factory(conv_module, provider(conv_module), "relu") + self.pool = bool(pool) + + def __call__(self, x: FeatureMap) -> FeatureMap: + import ttnn + + y = self.conv1(x) + if not self.pool: + return y + out = ttnn.max_pool2d(y.tensor, batch_size=y.batch, input_h=y.height, input_w=y.width, + channels=y.channels, kernel_size=[3, 3], stride=[2, 2], padding=[1, 1], + dilation=[1, 1], ceil_mode=False, memory_config=ttnn.DRAM_MEMORY_CONFIG, + output_layout=ttnn.TILE_LAYOUT) + y.deallocate() + oh, ow = maxpool_out_hw(y.height, y.width) + return FeatureMap(out, y.batch, oh, ow, y.channels) + + def convs(self) -> List[Any]: + return [self.conv1] + + +class ResNetStages: + """``len(layers)`` stages of ``block`` (``"bottleneck"`` or ``"basic"``); block 0 of stage s has the projection + when ``downsample_first[s]`` (default: every stage). ``__call__(x, keep)`` returns the outputs of the stages in + ``keep`` (others are freed once consumed; ``x`` itself is freed only with ``free_input=True``).""" + + def __init__(self, provider: Provider, names: Names, layers: Sequence[int], *, block: str = "bottleneck", + factory: Optional[ConvFactory] = None, downsample_first: Sequence[bool] = ()): + if block not in ("bottleneck", "basic"): + raise ValueError("block must be 'bottleneck' or 'basic'") + factory = factory or conv_factory() + cls = Bottleneck if block == "bottleneck" else BasicBlock + ds = list(downsample_first) or [True] * len(layers) + if len(ds) != len(layers): + raise ValueError("downsample_first needs one entry per stage") + self.block = block + self.stages: List[List[Any]] = [] + for s, n in enumerate(layers): + self.stages.append([cls(provider, names(s, b), has_downsample=(b == 0 and bool(ds[s])), factory=factory) + for b in range(int(n))]) + + def __call__(self, x: FeatureMap, keep: Sequence[int] = (), *, free_input: bool = False) -> List[FeatureMap]: + outs: List[FeatureMap] = [] + cur = x + for s, blocks in enumerate(self.stages): + for blk in blocks: + nxt = blk(cur) + if (cur is not x or free_input) and not any(cur is o for o in outs): + cur.deallocate() + cur = nxt + if s in keep: + outs.append(cur) + if cur is not x and not any(cur is o for o in outs): + cur.deallocate() + return outs + + def convs(self) -> List[Any]: + return [c for blocks in self.stages for blk in blocks for c in blk.convs()] + + +def resnet_numpy(x: np.ndarray, provider: Provider, names: Names, layers: Sequence[int], *, + block: str = "bottleneck", stem: Optional[str] = None, keep: Sequence[int] = (), + downsample_first: Sequence[bool] = ()) -> List[np.ndarray]: + """fp32 torch oracle of :class:`ResNetStem` (if ``stem`` names its conv) + :class:`ResNetStages` on an NCHW + array -> the kept stage outputs (NCHW float32).""" + import torch + import torch.nn.functional as F + + def conv(t, module, relu): + p = provider(module) + top, bottom, left, right = p.padding + t = F.pad(t, (left, right, top, bottom)) + y = F.conv2d(t, torch.from_numpy(np.asarray(p.weight, np.float32)), + None if p.bias is None else torch.from_numpy(np.asarray(p.bias, np.float32)), + stride=tuple(p.stride), dilation=tuple(p.dilation), groups=p.groups) + return torch.relu(y) if relu else y + + ds = list(downsample_first) or [True] * len(layers) + outs = [] + with torch.no_grad(): + t = torch.from_numpy(np.asarray(x, np.float32)) + if stem is not None: + t = F.max_pool2d(conv(t, stem, True), 3, 2, 1) + for s, n in enumerate(layers): + for b in range(int(n)): + nm = names(s, b) + idt = conv(t, nm["downsample"], False) if (b == 0 and ds[s]) else t + y = conv(t, nm["conv1"], True) + if block == "bottleneck": + y = conv(conv(y, nm["conv2"], True), nm["conv3"], False) + else: + y = conv(y, nm["conv2"], False) + t = torch.relu(y + idt) + if s in keep: + outs.append(t.numpy()) + return outs diff --git a/code/tt_diffusion_planner/ttaw/models/second.py b/code/tt_diffusion_planner/ttaw/models/second.py new file mode 100644 index 0000000000000000000000000000000000000000..3043a62778faa2925e4f44311e641f3ef7e738b5 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/models/second.py @@ -0,0 +1,311 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C24: SECOND backbone + SECONDFPN neck on the device (PLAN.md section 1.2 row C24; CenterPoint, PointPainting, +TransFusion and BEVFusion use them), built from BN-folded host weights on the C17 builders (``ttaw.ops.conv``). + +- :class:`Second`: blocks of 3x3 conv (zero padding 1) + ReLU; the first conv of a block carries the block stride + (mmdet3d ``SECOND``: ``layer_strides`` [1, 2, 2], CenterPoint-tiny [2, 2, 2]). Returns every block output. +- :class:`SecondFPN`: one deblock per block output, each up-sampling to the same grid, + ReLU (mmdet3d + ``SECONDFPN``): a ``"conv"`` deblock is a 1x1 stride-1 conv (a :class:`RowLinear`), a ``"conv_transpose"`` deblock + a ConvTranspose with kernel == stride (k = 1 a :class:`RowLinear`, k = 2 ``ttnn.conv_transpose2d``, k >= 3 + linear + depth-to-space: C17 ``ConvTranspose2d``, probe P5). The concat of the deblock outputs is NOT built: the + consumer takes the list (C25's K-split shared conv reads the parts, so CenterPoint's 177 MB concat never exists). + +Placement: every module output is a DRAM interleaved TILE ``[1, 1, N*H*W, C]`` map (the robust first-port layout: +3x3 convs on DRAM inputs run ttnn's automatic DRAM slicing, probe P14; L1 residency is optimization work). +Precision: :data:`DEFAULT_PRECISION` (C17's ``CONV_PRECISION``, HiFi4 + fp32 accumulation + packer L1 +accumulation: without it a conv keeps its partial sums in bf16 between inner blocks) unless a +``ttaw.precision.PrecisionPolicy`` says otherwise for ``.block.conv`` / ``.deblock``; a +precision's ``activations`` dtype (``:a=fp32``) is the dtype of that layer's OUTPUT (bf16 by default), the knob of a +deep plain conv stack whose bf16 activation rounding compounds (S:centerpoint:396-408). + +:class:`ConvSpec` records are plain numpy, so a port maps its own weight loader onto them; :func:`second_numpy` / +:func:`second_fpn_numpy` are the fp32 torch oracles of the device modules. +""" +from __future__ import annotations + +from dataclasses import dataclass +from typing import Any, Dict, List, Optional, Sequence + +import numpy as np + +from ..ops.conv import CONV_PRECISION, Conv2d, ConvTranspose2d, FeatureMap, is_dram +from ..precision import Precision, PrecisionPolicy + +__all__ = ["ConvSpec", "RowLinear", "Second", "SecondFPN", "DEFAULT_PRECISION", "default_policy", "second_numpy", + "second_fpn_numpy", "free_maps"] + +DEFAULT_PRECISION = CONV_PRECISION + + +@dataclass(frozen=True) +class ConvSpec: + """One conv layer with BN folded: ``y = relu?(conv(x, weight) + bias)``. + + ``kind="conv"``: weight ``[Cout, Cin, k, k]``; ``kind="conv_transpose"``: weight ``[Cin, Cout, s, s]`` (kernel == + stride, no padding). ``padding`` None = ``k // 2`` for convs.""" + + weight: np.ndarray + bias: Optional[np.ndarray] + stride: int = 1 + padding: Optional[int] = None + relu: bool = True + kind: str = "conv" + name: str = "" + + @property + def kernel(self) -> int: + return int(np.shape(self.weight)[-1]) + + @property + def in_channels(self) -> int: + return int(np.shape(self.weight)[1 if self.kind == "conv" else 0]) + + @property + def out_channels(self) -> int: + return int(np.shape(self.weight)[0 if self.kind == "conv" else 1]) + + +def default_policy() -> PrecisionPolicy: + """:data:`DEFAULT_PRECISION` everywhere (bf16 weights).""" + return PrecisionPolicy(default=DEFAULT_PRECISION) + + +def free_maps(*maps: Any) -> None: + """Deallocate device feature maps / tensors that still hold their buffer (``force=False``: memory shared with + a view or another owner is left alone; ``ttnn.deallocate`` forces by default, probe P2).""" + import ttnn + + for m in maps: + t = m.tensor if isinstance(m, FeatureMap) else m + if t is not None and t.is_allocated(): + ttnn.deallocate(t, False) + + +def _to_dram(fm: FeatureMap) -> FeatureMap: + """The map in DRAM (moved to DRAM interleaved when an op left it in L1, e.g. ``conv_transpose2d``'s sharded + output; a no-op for DRAM maps).""" + import ttnn + + if is_dram(fm.tensor): + return fm + out = fm.with_tensor(ttnn.to_memory_config(fm.tensor, ttnn.DRAM_MEMORY_CONFIG)) + free_maps(fm) + return out + + +class RowLinear: + """``y = act(x @ W + b)`` on the rows of a feature map (a 1x1 stride-1 conv): ``[1, 1, R, K]`` TILE -> + ``[1, 1, R, N]`` TILE in DRAM, through ``ttnn.linear`` with an explicit compute config and the device grid as + ``core_grid`` (which keeps the activation fused in the matmul). ``weight``: ``[K, N]`` (in, out) or a 1x1 conv + weight ``[N, K, 1, 1]``. Device weights are uploaded by the first call (outside any capture).""" + + def __init__(self, weight: Any, bias: Any = None, *, activation: Optional[str] = None, + precision: Any = DEFAULT_PRECISION, output_dtype: str = "bfloat16", name: str = "linear"): + w = np.asarray(weight, dtype=np.float32) + if w.ndim == 4: + if w.shape[2:] != (1, 1): + raise ValueError(f"{name}: a conv weight must be 1x1, got {w.shape}") + w = w[:, :, 0, 0].T + if w.ndim != 2: + raise ValueError(f"{name}: weight must be [K, N], got {w.shape}") + if activation not in (None, "relu"): + raise ValueError(f"{name}: activation {activation!r}: None or 'relu'") + self.name = name + self.weight_host = np.ascontiguousarray(w) + self.in_channels, self.out_channels = (int(s) for s in w.shape) + b = np.zeros(self.out_channels, np.float32) if bias is None else np.asarray(bias, np.float32) + if b.shape != (self.out_channels,): + raise ValueError(f"{name}: bias must be ({self.out_channels},), got {b.shape}") + if not (np.all(np.isfinite(w)) and np.all(np.isfinite(b))): + raise ValueError(f"{name}: non-finite weights") + self.bias_host = b + self.activation = activation + self.precision = Precision.parse(precision) + self.output_dtype = output_dtype + self._device: Optional[tuple] = None + + def _upload(self, device: Any): + import torch + import ttnn + + from ..tensors import ttnn_dtype + + wdt = ttnn_dtype(self.precision.weights) + bdt = ttnn.bfloat16 if wdt == ttnn.bfloat8_b else wdt + g = device.compute_with_storage_grid_size() + self._device = ( + ttnn.from_torch(torch.from_numpy(self.weight_host), dtype=wdt, layout=ttnn.TILE_LAYOUT, device=device, + memory_config=ttnn.DRAM_MEMORY_CONFIG), + ttnn.from_torch(torch.from_numpy(self.bias_host.reshape(1, -1)), dtype=bdt, layout=ttnn.TILE_LAYOUT, + device=device, memory_config=ttnn.DRAM_MEMORY_CONFIG), + ttnn.CoreGrid(y=int(g.y), x=int(g.x)), + self.precision.compute_kernel_config()) + return self._device + + def __call__(self, x: Any) -> Any: + """``x``: a TILE device tensor ``[..., R, K]`` or a :class:`FeatureMap` (returned as one).""" + import ttnn + + from ..tensors import ttnn_dtype + + t = x.tensor if isinstance(x, FeatureMap) else x + if int(t.shape[-1]) != self.in_channels: + raise ValueError(f"{self.name}: input has {int(t.shape[-1])} channels, expected {self.in_channels}") + weight, bias, grid, cfg = self._device or self._upload(t.device()) + y = ttnn.linear(t, weight, bias=bias, activation=self.activation, compute_kernel_config=cfg, core_grid=grid, + memory_config=ttnn.DRAM_MEMORY_CONFIG, dtype=ttnn_dtype(self.output_dtype)) + return x.with_tensor(y, self.out_channels) if isinstance(x, FeatureMap) else y + + def release(self) -> None: + if self._device is not None: + free_maps(self._device[0], self._device[1]) + self._device = None + + def describe(self) -> Dict[str, Any]: + return {"name": self.name, "cin": self.in_channels, "cout": self.out_channels, "activation": self.activation, + "precision": self.precision.label, "output_dtype": self.output_dtype} + + +def _precision(policy: Optional[PrecisionPolicy], name: str) -> Precision: + return policy.resolve(name) if policy is not None else Precision.parse(DEFAULT_PRECISION) + + +class Second: + """SECOND backbone: ``blocks[i]`` is the list of :class:`ConvSpec` of block ``i`` (k x k convs, BN folded).""" + + def __init__(self, blocks: Sequence[Sequence[ConvSpec]], *, policy: Optional[PrecisionPolicy] = None, + prefix: str = "backbone"): + self.prefix = prefix + self.blocks: List[List[Conv2d]] = [] + for i, block in enumerate(blocks): + layers = [] + for j, spec in enumerate(block): + if spec.kind != "conv": + raise ValueError(f"{prefix}.block{i}.conv{j}: SECOND layers are convs, got {spec.kind}") + name = f"{prefix}.block{i}.conv{j}" + pad = spec.kernel // 2 if spec.padding is None else int(spec.padding) + prec = _precision(policy, name) + layers.append(Conv2d(spec.weight, spec.bias, stride=spec.stride, padding=pad, + activation="relu" if spec.relu else None, precision=prec, + output_dtype=prec.activations, name=name)) + self.blocks.append(layers) + + def out_shapes(self, height: int, width: int) -> List[tuple]: + """``(H, W, C)`` of every block output for an input of ``height x width``.""" + shapes = [] + for block in self.blocks: + for layer in block: + height, width = layer.out_hw(height, width) + shapes.append((height, width, block[-1].out_channels)) + return shapes + + def __call__(self, x: FeatureMap) -> List[FeatureMap]: + """Run every block on ``x`` (kept); returns the block outputs (DRAM). Inner layer outputs are freed.""" + outs: List[FeatureMap] = [] + cur = x + for block in self.blocks: + for layer in block: + nxt = _to_dram(layer(cur)) + if cur is not x and all(cur is not o for o in outs): + free_maps(cur) + cur = nxt + outs.append(cur) + return outs + + def release(self) -> None: + for block in self.blocks: + for layer in block: + layer.release() + + def describe(self) -> List[List[Dict[str, Any]]]: + return [[layer.describe() for layer in block] for block in self.blocks] + + +class SecondFPN: + """SECONDFPN neck: ``deblocks[i]`` (a :class:`ConvSpec`, ReLU after) up-samples block output ``i``.""" + + def __init__(self, deblocks: Sequence[ConvSpec], *, policy: Optional[PrecisionPolicy] = None, + prefix: str = "neck"): + self.prefix = prefix + self.deblocks: List[Any] = [] + self.strides: List[int] = [] + for i, spec in enumerate(deblocks): + name = f"{prefix}.deblock{i}" + act = "relu" if spec.relu else None + prec = _precision(policy, name) + if spec.kind == "conv": + if spec.kernel != 1 or spec.stride != 1: + raise ValueError(f"{name}: a conv deblock must be 1x1 stride 1 (k={spec.kernel}, s={spec.stride})") + layer: Any = RowLinear(spec.weight, spec.bias, activation=act, precision=prec, + output_dtype=prec.activations, name=name) + self.strides.append(1) + elif spec.kind == "conv_transpose": + if spec.kernel != spec.stride: + raise ValueError(f"{name}: ConvTranspose deblocks need kernel == stride (k={spec.kernel}, " + f"s={spec.stride})") + if spec.stride == 1: + layer = RowLinear(np.asarray(spec.weight, np.float32)[:, :, 0, 0], spec.bias, activation=act, + precision=prec, output_dtype=prec.activations, name=name) + else: + layer = ConvTranspose2d(spec.weight, spec.bias, stride=spec.stride, activation=act, + precision=prec, output_dtype=prec.activations, name=name) + self.strides.append(spec.stride) + else: + raise ValueError(f"{name}: unknown kind {spec.kind!r}") + self.deblocks.append(layer) + + def __call__(self, blocks: Sequence[FeatureMap]) -> List[FeatureMap]: + """The deblock outputs (DRAM interleaved TILE maps of one grid); the inputs are kept.""" + if len(blocks) != len(self.deblocks): + raise ValueError(f"{self.prefix}: {len(blocks)} inputs for {len(self.deblocks)} deblocks") + outs = [] + for layer, x in zip(self.deblocks, blocks): + outs.append(_to_dram(layer(x))) + return outs + + def release(self) -> None: + for d in self.deblocks: + d.release() + + def describe(self) -> List[Dict[str, Any]]: + return [d.describe() for d in self.deblocks] + + +# ------------------------------------------------------------------------------------------- host oracles + +def _apply_torch(spec: ConvSpec, x): + """One layer in torch fp32 on an NCHW tensor.""" + import torch + import torch.nn.functional as F + + w = torch.from_numpy(np.ascontiguousarray(spec.weight, dtype=np.float32)) + b = None if spec.bias is None else torch.from_numpy(np.ascontiguousarray(spec.bias, dtype=np.float32)) + if spec.kind == "conv": + pad = spec.kernel // 2 if spec.padding is None else spec.padding + y = F.conv2d(x, w, b, stride=spec.stride, padding=pad) + else: + y = F.conv_transpose2d(x, w, b, stride=spec.stride) + return torch.relu(y) if spec.relu else y + + +def second_numpy(blocks: Sequence[Sequence[ConvSpec]], x_nchw: Any) -> List[np.ndarray]: + """fp32 oracle of :class:`Second` (torch): NCHW in, NCHW block outputs out.""" + import torch + + with torch.no_grad(): + x = torch.as_tensor(np.asarray(x_nchw, dtype=np.float32)) + outs = [] + for block in blocks: + for spec in block: + x = _apply_torch(spec, x) + outs.append(x.numpy().copy()) + return outs + + +def second_fpn_numpy(deblocks: Sequence[ConvSpec], blocks_nchw: Sequence[Any]) -> List[np.ndarray]: + """fp32 oracle of :class:`SecondFPN` (torch): NCHW in and out.""" + import torch + + with torch.no_grad(): + return [_apply_torch(s, torch.as_tensor(np.asarray(x, dtype=np.float32))).numpy().copy() + for s, x in zip(deblocks, blocks_nchw)] diff --git a/code/tt_diffusion_planner/ttaw/models/transfusion_head.py b/code/tt_diffusion_planner/ttaw/models/transfusion_head.py new file mode 100644 index 0000000000000000000000000000000000000000..1e7010f07ca83b5b39f9ff7159de3935a257c759 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/models/transfusion_head.py @@ -0,0 +1,584 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C26: the TransFusion query head on the device (PLAN.md section 1.2 row C26; TransFusion builds it, BEVFusion and +PTv3 reuse it; probes P4, P7, P8, P13). + +The dense part of a TransFusion head (shared conv, heatmap convs) is C17 / C24 work; this module is what follows the +local-max heatmap (C22): query selection (C21), query initialisation, ONE post-norm decoder layer and the prediction +heads, as one traceable device graph with one packed output. The model-agnostic weights record +:class:`QueryHeadWeights` (host numpy, BN folded) describes it: + +- ``class_table`` ``(C, d)``: the class encoding (``Conv1d(one_hot(c))`` = ``W[:, c] + b``, an exact row gather); +- ``bev_pos`` ``(cells, 2)``: the model's constant BEV position of every feature cell (TF / PT ``(x + 0.5, y + 0.5)`` + with ``k = y * W + x``; BF ``(i + 0.5, j + 0.5)`` with ``k = i * 180 + j``); a query's position is ``bev_pos[pos]``; +- ``self_posembed`` / ``cross_posembed``: the position MLPs ``Conv1d 2 -> d + ReLU, Conv1d d -> d`` (BN folded); +- the decoder layer: self-attention on ``s = x + pe(query_pos)`` (Q = K = V = s), add + LayerNorm, cross-attention + of ``x + pe(query_pos)`` over ``key = lidar_feat + pe_k(bev_pos)`` (K = V = key), add + LN, FFN (ReLU), add + LN; +- the prediction heads ``(name, Conv1d d -> h + ReLU, Conv1d h -> c)``. + +Device form (every rewrite exact in real arithmetic; the oracle :class:`QueryHeadOracle` runs the literal forms): + +| step | device ops | +|---|---| +| selection | C21 ``TopKSelect``: class-major bf16 row of the local-max heatmap, ``topk_large_indices`` k = ``padded_k(K)``, exact fp32 decode -> uint32 ``pos`` / ``cls`` rows (or given indices: teacher forcing) | +| query init | ``ttnn.embedding`` of the ``lidar_feat`` rows (bf16) by ``pos`` + the class-table row by ``cls``; the query position embedding is a function of ``pos`` only, so it is a constant **QPE table** ``self_posembed(bev_pos)`` (host float64 -> bf16) gathered by ``pos``: ``query_pos`` itself never passes through bf16 (S:transfusion:241) | +| keys | ``key = lidar_feat + KPE`` with the constant ``KPE = cross_posembed(bev_pos)`` table (the 0.6 GMAC position branch leaves the per-frame graph, S:transfusion:279) | +| attention | C20 ``sdpa`` (``is_causal=False``, ``scale`` from the weights, chunk table), the head dim zero-padded to 32 in the projection weights (``pad_head_columns`` / ``pad_head_rows``, exact); self-attention on the logical ``K`` queries with no mask, Q / K / V from one fused projection; cross-attention over every cell | +| LayerNorm | ``ttnn.layer_norm(o, residual_input_tensor=x, epsilon=)`` (never ttnn's 1e-12 default) | +| heads | merged: ``relu(x @ [Wh_1 ... Wh_n] + bh)`` then the block-diagonal ``@ Wo + bo`` (padded to 32), fp32 out | +| query heat | ``ttnn.embedding`` of the local-max heatmap rows (classes zero-padded to 32) by ``pos``: every class at the query's cell (the TF / mmdeploy score ``query_heat * sigmoid(heatmap)``; BF / PT one-hot it on the host) | + +Output (``__call__``): ``{"heads": fp32 [1, 1, K, 32], "query_heat": bf16 [1, 1, Kp, 32], "indices": fp32 [1, 1, 1, +Kp]}`` (``Kp = padded_k(K)``; the first ``K`` indices are the proposals, fp32 holds them exactly), ready for +``ttaw.trace.pack_outputs``. :func:`assemble_transfusion` turns the read-back arrays into the TF / mmdeploy graph +outputs ``cls_score0``, ``bbox_pred0``, ``dir_cls_pred0`` (``center + query_pos`` in float32 on the host). + +Limits: one decoder layer (the TF, BF and PT heads have one; a second layer would re-embed the predicted centres, +which no constant table can do); ``bev_pos`` must be finite. Precision (``ttaw.precision`` policy, module names +``.self_attn``, ``.cross_attn``, ``.ffn``, ``.heads``, ``.norm``): fp32 weights, HiFi4 + fp32 accumulation, +bf16 activations (SDPA takes bf16); the tables are bf16 (their rows are model activations, rounded once). + +numpy only at import; ttnn / torch are imported inside the functions that need them. +""" +from __future__ import annotations + +from dataclasses import dataclass +from typing import Any, Dict, List, Optional, Sequence, Tuple + +import numpy as np + +from ..ops.attention import pad_head_columns, pad_head_rows, padded_head_dim +from ..ops.topk import TopKSelect, padded_k +from ..precision import Precision, PrecisionPolicy +from .second import RowLinear, free_maps + +__all__ = ["MHAWeights", "PosEmbedMLP", "PredictionHead", "QueryHeadWeights", "TransFusionQueryHead", + "QueryHeadOracle", "assemble_transfusion", "DEFAULT_PRECISION", "default_policy", "HEAD_PAD"] + +HEAD_PAD = 32 +DEFAULT_PRECISION = "HiFi4+fp32:w=fp32" + + +def default_policy() -> PrecisionPolicy: + """fp32 weights, HiFi4 + fp32 accumulation, bf16 activations for every module of the head.""" + return PrecisionPolicy(default=DEFAULT_PRECISION) + + +def _f32(a: Any) -> np.ndarray: + """A writable float32 copy (ONNX-derived arrays may be read-only; torch warns about those).""" + return np.array(a, dtype=np.float32, copy=True, order="C") + + +# ------------------------------------------------------------------------------------------------- weights + +@dataclass(frozen=True) +class MHAWeights: + """``nn.MultiheadAttention`` as exported, ``[in, out]`` layout: ``q = x_q @ wq + bq`` (likewise k / v), per head + ``softmax(scale * q k^T) v``, ``out = concat_heads @ wo + bo``.""" + + wq: np.ndarray + bq: np.ndarray + wk: np.ndarray + bk: np.ndarray + wv: np.ndarray + bv: np.ndarray + wo: np.ndarray + bo: np.ndarray + num_heads: int + scale: float + + @property + def head_dim(self) -> int: + return int(np.shape(self.wq)[1]) // int(self.num_heads) + + @property + def dim(self) -> int: + return int(np.shape(self.wq)[0]) + + +@dataclass(frozen=True) +class PosEmbedMLP: + """``Conv1d 2 -> d`` (BN folded) + ReLU, ``Conv1d d -> d``; ``[in, out]`` weights.""" + + w0: np.ndarray + b0: np.ndarray + w1: np.ndarray + b1: np.ndarray + + def numpy(self, pos: Any) -> np.ndarray: + """``(P, 2)`` positions -> ``(P, d)`` float64.""" + p = np.asarray(pos, np.float64) + h = np.maximum(p @ np.asarray(self.w0, np.float64) + np.asarray(self.b0, np.float64), 0.0) + return h @ np.asarray(self.w1, np.float64) + np.asarray(self.b1, np.float64) + + +@dataclass(frozen=True) +class PredictionHead: + """``Conv1d d -> h`` + ReLU, ``Conv1d h -> c`` (``[in, out]`` weights).""" + + name: str + w_hidden: np.ndarray + b_hidden: np.ndarray + w_out: np.ndarray + b_out: np.ndarray + + @property + def channels(self) -> int: + return int(np.shape(self.w_out)[1]) + + +@dataclass(frozen=True) +class QueryHeadWeights: + """Everything the query head needs, model-agnostic (module docstring). ``height`` / ``width``: the feature grid + (``cells = height * width`` keys, row-major ``k = y * width + x`` as the dense maps are laid out).""" + + height: int + width: int + num_classes: int + num_proposals: int + class_table: np.ndarray # (C, d) + bev_pos: np.ndarray # (cells, 2) + self_posembed: PosEmbedMLP + cross_posembed: PosEmbedMLP + self_attn: MHAWeights + cross_attn: MHAWeights + norms: Tuple[Tuple[np.ndarray, np.ndarray, float], ...] # (gamma, beta, eps) x 3: after self, cross, FFN + ffn: Tuple[np.ndarray, np.ndarray, np.ndarray, np.ndarray] # w1 (d, f), b1, w2 (f, d), b2 + heads: Tuple[PredictionHead, ...] + + @property + def cells(self) -> int: + return int(self.height) * int(self.width) + + @property + def dim(self) -> int: + return int(np.shape(self.class_table)[1]) + + def validate(self) -> None: + d, c, cells = self.dim, int(self.num_classes), self.cells + if np.shape(self.class_table) != (c, d): + raise ValueError(f"class_table {np.shape(self.class_table)} != ({c}, {d})") + if np.shape(self.bev_pos) != (cells, 2) or not np.all(np.isfinite(self.bev_pos)): + raise ValueError(f"bev_pos must be finite ({cells}, 2), got {np.shape(self.bev_pos)}") + for a in (self.self_attn, self.cross_attn): + if a.dim != d or np.shape(a.wo) != (a.num_heads * a.head_dim, d): + raise ValueError("attention weights do not match the model width") + if not (np.isfinite(a.scale) and a.scale > 0): + raise ValueError(f"attention scale {a.scale}") + if len(self.norms) != 3: + raise ValueError("a decoder layer has three LayerNorms") + if np.shape(self.ffn[0])[0] != d or np.shape(self.ffn[2])[1] != d: + raise ValueError("FFN weights do not match the model width") + if not self.heads or any(np.shape(h.w_hidden)[0] != d for h in self.heads): + raise ValueError("prediction heads do not read the model width") + if sum(h.channels for h in self.heads) > HEAD_PAD: + raise ValueError(f"the merged head output is padded to {HEAD_PAD} channels") + if c > HEAD_PAD: + raise ValueError(f"the query heat rows hold at most {HEAD_PAD} classes") + padded_k(self.num_proposals) + if self.num_proposals > cells * c: + raise ValueError("more proposals than heatmap entries") + + # ---- host tables (float64, rounded once by the device upload) ------------------------------------------ + def query_pos_table(self) -> np.ndarray: + """``self_posembed(bev_pos)`` ``(cells, d)``: the query position embedding of every cell.""" + return self.self_posembed.numpy(self.bev_pos) + + def key_pos_table(self) -> np.ndarray: + """``cross_posembed(bev_pos)`` ``(cells, d)``: the input-independent key position embedding.""" + return self.cross_posembed.numpy(self.bev_pos) + + def merged_heads(self, pad_to: int = HEAD_PAD): + """``(Wh, bh, Wo, bo, slices)``: ``relu(x @ Wh + bh) @ Wo + bo`` with the block-diagonal ``Wo`` padded to + ``pad_to`` columns; ``slices`` = ``[(name, start, stop)]`` of each head's output channels (exact rewrite).""" + wh = np.concatenate([_f32(h.w_hidden) for h in self.heads], axis=1) + bh = np.concatenate([_f32(h.b_hidden) for h in self.heads]) + rows = wh.shape[1] + cols = sum(h.channels for h in self.heads) + width = max(pad_to, -(-cols // pad_to) * pad_to) + wo = np.zeros((rows, width), np.float32) + bo = np.zeros(width, np.float32) + slices, r, c = [], 0, 0 + for h in self.heads: + hr, hc = np.shape(h.w_hidden)[1], h.channels + wo[r:r + hr, c:c + hc] = _f32(h.w_out) + bo[c:c + hc] = _f32(h.b_out) + slices.append((h.name, c, c + hc)) + r, c = r + hr, c + hc + return wh, bh, wo, bo, slices + + +# ---------------------------------------------------------------------------------------------- device + +class _LayerNorm: + def __init__(self, gamma: Any, beta: Any, eps: float, precision: Precision, name: str): + self.gamma, self.beta, self.eps = _f32(gamma).reshape(1, -1).copy(), _f32(beta).reshape(1, -1).copy(), float(eps) + self.precision = precision + self.name = name + self._dev: Optional[Tuple[Any, Any, Any]] = None + + def upload(self, device: Any): + import torch + import ttnn + + def up(a): + return ttnn.from_torch(torch.from_numpy(a), dtype=ttnn.float32, layout=ttnn.TILE_LAYOUT, device=device, + memory_config=ttnn.DRAM_MEMORY_CONFIG) + + self._dev = (up(self.gamma), up(self.beta), self.precision.compute_kernel_config()) + return self._dev + + def __call__(self, x: Any, residual: Any): + """``LayerNorm(x + residual)`` (one ``ttnn.layer_norm`` with a fused residual add).""" + import ttnn + + g, b, cfg = self._dev or self.upload(x.device()) + return ttnn.layer_norm(x, residual_input_tensor=residual, weight=g, bias=b, epsilon=self.eps, + compute_kernel_config=cfg, memory_config=ttnn.DRAM_MEMORY_CONFIG) + + def release(self) -> None: + if self._dev is not None: + free_maps(self._dev[0], self._dev[1]) + self._dev = None + + +class TransFusionQueryHead: + """The query head of :class:`QueryHeadWeights` on the device (module docstring). Device weights and tables are + uploaded by :meth:`prepare` (or the first call), outside any capture; the forward is traceable.""" + + def __init__(self, weights: QueryHeadWeights, *, policy: Optional[PrecisionPolicy] = None, prefix: str = "head", + attn_fp32_acc: bool = False, name: str = "query_head"): + weights.validate() + self.weights = weights + self.prefix = prefix + self.name = name + self.policy = policy or default_policy() + self.attn_fp32_acc = bool(attn_fp32_acc) + self.dim = weights.dim + self.cells = weights.cells + self.classes = int(weights.num_classes) + self.k = int(weights.num_proposals) + self.select = TopKSelect(self.cells, self.classes, self.k, name=f"{prefix}.topk") + self.kp = self.select.kp + sa, ca = weights.self_attn, weights.cross_attn + self.heads_self, self.hd_self = int(sa.num_heads), sa.head_dim + self.heads_cross, self.hd_cross = int(ca.num_heads), ca.head_dim + self.scale_self, self.scale_cross = float(sa.scale), float(ca.scale) + + def prec(module: str) -> Precision: + return self.policy.resolve(f"{prefix}.{module}") + + def cols(w, b, a: MHAWeights): + return pad_head_columns(_f32(w), _f32(b), a.num_heads, a.head_dim) + + p_sa, p_ca, p_ffn, p_heads = prec("self_attn"), prec("cross_attn"), prec("ffn"), prec("heads") + q, bq = cols(sa.wq, sa.bq, sa) + k, bk = cols(sa.wk, sa.bk, sa) + v, bv = cols(sa.wv, sa.bv, sa) + self.self_qkv = RowLinear(np.concatenate([q, k, v], axis=1), np.concatenate([bq, bk, bv]), precision=p_sa, + name=f"{prefix}.self_attn.qkv") + self.self_out = RowLinear(pad_head_rows(_f32(sa.wo), sa.num_heads, sa.head_dim), _f32(sa.bo), + precision=p_sa, name=f"{prefix}.self_attn.out") + self.cross_q = RowLinear(*cols(ca.wq, ca.bq, ca), precision=p_ca, name=f"{prefix}.cross_attn.q") + self.cross_k = RowLinear(*cols(ca.wk, ca.bk, ca), precision=p_ca, name=f"{prefix}.cross_attn.k") + self.cross_v = RowLinear(*cols(ca.wv, ca.bv, ca), precision=p_ca, name=f"{prefix}.cross_attn.v") + self.cross_out = RowLinear(pad_head_rows(_f32(ca.wo), ca.num_heads, ca.head_dim), _f32(ca.bo), + precision=p_ca, name=f"{prefix}.cross_attn.out") + w1, b1, w2, b2 = weights.ffn + self.ffn1 = RowLinear(_f32(w1), _f32(b1), activation="relu", precision=p_ffn, name=f"{prefix}.ffn.0") + self.ffn2 = RowLinear(_f32(w2), _f32(b2), precision=p_ffn, name=f"{prefix}.ffn.1") + wh, bh, wo, bo, self.head_slices = weights.merged_heads() + self.heads_hidden = RowLinear(wh, bh, activation="relu", precision=p_heads, name=f"{prefix}.heads.hidden") + self.heads_out = RowLinear(wo, bo, precision=p_heads, output_dtype="float32", name=f"{prefix}.heads.out") + p_norm = prec("norm") + self.norms = [_LayerNorm(g, b, eps, p_norm, f"{prefix}.norm{i}") for i, (g, b, eps) in enumerate(weights.norms)] + self._tables: Optional[Dict[str, Any]] = None + + # ---- device constants -------------------------------------------------------------------------------- + def prepare(self, device: Any) -> None: + """Upload the tables (bf16: class table and QPE as ROW_MAJOR gather tables, KPE as a TILE map) and the + LayerNorm rows. The linears upload their weights on their first call (the warm-up).""" + import torch + import ttnn + + if self._tables is not None: + return + w = self.weights + + def up(a: np.ndarray, layout): + t = torch.from_numpy(_f32(a).reshape(1, 1, a.shape[0], a.shape[1])) + return ttnn.from_torch(t, dtype=ttnn.bfloat16, layout=layout, device=device, + memory_config=ttnn.DRAM_MEMORY_CONFIG) + + self._tables = {"class": up(_f32(w.class_table), ttnn.ROW_MAJOR_LAYOUT), + "qpe": up(w.query_pos_table(), ttnn.ROW_MAJOR_LAYOUT), + "kpe": up(w.key_pos_table(), ttnn.TILE_LAYOUT)} + for n in self.norms: + n.upload(device) + + def _t(self, device: Any) -> Dict[str, Any]: + if self._tables is None: + self.prepare(device) + return self._tables # type: ignore[return-value] + + # ---- forward pieces ------------------------------------------------------------------------------------ + @staticmethod + def _rows(t: Any, n: int): + """``[1, n, D]`` / ``[1, 1, n, D]`` -> ``[1, 1, n, D]`` (a view).""" + import ttnn + + return ttnn.reshape(t, (1, 1, int(n), int(t.shape[-1]))) + + def _gather(self, index: Any, table: Any): + import ttnn + + return self._rows(ttnn.embedding(index, table, layout=ttnn.TILE_LAYOUT), int(index.shape[-1])) + + def init_queries(self, lidar_feat: Any, pos: Any, cls: Any) -> Tuple[Any, Any]: + """``lidar_feat`` TILE ``[1, 1, cells, d]``, ``pos`` / ``cls`` uint32 ROW_MAJOR ``[1, Kp]`` -> ``(x, qpe)`` + TILE ``[1, 1, K, d]``: the initial queries ``lidar_feat[pos] + class_table[cls]`` and their position + embeddings ``QPE[pos]``.""" + import ttnn + + t = self._t(lidar_feat.device()) + table = ttnn.to_layout(lidar_feat, ttnn.ROW_MAJOR_LAYOUT) + feat = self._gather(pos, table) + free_maps(table) + cemb = self._gather(cls, t["class"]) + x_p = ttnn.add(feat, cemb) + free_maps(feat, cemb) + qpe_p = self._gather(pos, t["qpe"]) + if self.kp == self.k: + return x_p, qpe_p + x = ttnn.slice(x_p, [0, 0, 0, 0], [1, 1, self.k, self.dim]) + qpe = ttnn.slice(qpe_p, [0, 0, 0, 0], [1, 1, self.k, self.dim]) + free_maps(x_p, qpe_p) + return x, qpe + + def keys(self, lidar_feat: Any): + """``lidar_feat + KPE`` (TILE ``[1, 1, cells, d]``): the cross-attention key / value input.""" + import ttnn + + return ttnn.add(lidar_feat, self._t(lidar_feat.device())["kpe"]) + + def _sdpa(self, q, k, v, scale: float): + from ..ops.attention import sdpa + + return sdpa(q, k, v, scale=scale, fp32_acc=self.attn_fp32_acc, concat_heads=True) + + def decoder(self, x: Any, qpe: Any, key: Any, taps: Optional[Dict[str, Any]] = None): + """One post-norm decoder layer on ``x`` ``[1, 1, K, d]`` -> ``[1, 1, K, d]`` (``key``: :meth:`keys`).""" + import ttnn + + from ..ops.attention import split_heads, split_qkv + + s = ttnn.add(x, qpe) + qkv = self.self_qkv(s) + free_maps(s) + q, k, v = split_qkv(qkv, self.heads_self) + free_maps(qkv) + a = self._sdpa(q, k, v, self.scale_self) + free_maps(q, k, v) + o = self.self_out(a) + free_maps(a) + x1 = self.norms[0](o, x) + free_maps(o) + if taps is not None: + taps["decoder.self_attn"] = x1 + qc_in = ttnn.add(x1, qpe) + qc = self.cross_q(qc_in) + free_maps(qc_in) + kc, vc = self.cross_k(key), self.cross_v(key) + qh, kh, vh = (split_heads(t, self.heads_cross) for t in (qc, kc, vc)) + free_maps(qc, kc, vc) + a = self._sdpa(qh, kh, vh, self.scale_cross) + free_maps(qh, kh, vh) + o = self.cross_out(a) + free_maps(a) + x2 = self.norms[1](o, x1) + free_maps(o) + if taps is not None: + taps["decoder.cross_attn"] = x2 + else: + free_maps(x1) + h = self.ffn1(x2) + f = self.ffn2(h) + free_maps(h) + x3 = self.norms[2](f, x2) + free_maps(f) + if taps is None: + free_maps(x2) + return x3 + + def predict(self, x: Any): + """The merged prediction heads: ``[1, 1, K, d]`` -> fp32 ``[1, 1, K, 32]`` (``head_slices`` channels).""" + hid = self.heads_hidden(x) + out = self.heads_out(hid) + free_maps(hid) + return out + + def query_heat(self, heat_nms: Any, pos: Any): + """The local-max heatmap rows (every class, zero-padded to 32) at the proposals' cells: bf16 TILE + ``[1, 1, Kp, 32]``.""" + import ttnn + + padded = ttnn.pad(heat_nms, [(0, 0), (0, 0), (0, 0), (0, HEAD_PAD - self.classes)], 0.0) + table = ttnn.to_layout(padded, ttnn.ROW_MAJOR_LAYOUT) + free_maps(padded) + out = self._gather(pos, table) + free_maps(table) + return out + + def _indices_tile(self, indices: Any): + """uint32 ROW_MAJOR ``[1, Kp]`` -> fp32 ``[1, 1, 1, Kp]`` ROW_MAJOR-ready tensor (exact below 2**24).""" + import ttnn + + t = ttnn.to_layout(indices, ttnn.TILE_LAYOUT) + f = ttnn.typecast(t, ttnn.float32) + free_maps(t) + return ttnn.reshape(f, (1, 1, 1, int(indices.shape[-1]))) + + def __call__(self, lidar_feat: Any, heat_nms: Any, indices: Optional[Any] = None, + taps: Optional[Dict[str, Any]] = None) -> Dict[str, Any]: + """``lidar_feat`` TILE ``[1, 1, cells, d]`` (bf16), ``heat_nms`` TILE ``[1, 1, cells, C]`` bf16 (C22's + output), ``indices``: None (device top-k) or a uint32 ROW_MAJOR ``[1, Kp]`` input (teacher forcing). Returns + ``{"heads", "query_heat", "indices"}`` (module docstring); ``taps`` (a dict) also collects ``query.feat``, + ``decoder.query_pos_embed``, ``decoder.self_attn``, ``decoder.cross_attn``, ``decoder.out``.""" + idx = self.select.indices(heat_nms) if indices is None else indices + pos, cls = self.select.decode(idx) + x, qpe = self.init_queries(lidar_feat, pos, cls) + free_maps(cls) + key = self.keys(lidar_feat) + y = self.decoder(x, qpe, key, taps) + free_maps(key) + heads = self.predict(y) + qheat = self.query_heat(heat_nms, pos) + free_maps(pos) + idx_f = self._indices_tile(idx) + if indices is None: + free_maps(idx) + if taps is not None: + taps.update({"query.feat": x, "decoder.query_pos_embed": qpe, "decoder.out": y}) + else: + free_maps(x, qpe, y) + return {"heads": heads, "query_heat": qheat, "indices": idx_f} + + # ---- host side ----------------------------------------------------------------------------------------- + def unpack_heads(self, heads: Any) -> Dict[str, np.ndarray]: + """The read-back ``heads`` array -> ``{name: (c, K)}`` float32.""" + a = np.asarray(heads, np.float32).reshape(-1, HEAD_PAD)[: self.k] + return {name: np.ascontiguousarray(a[:, s:e].T) for name, s, e in self.head_slices} + + def release(self) -> None: + for m in (self.self_qkv, self.self_out, self.cross_q, self.cross_k, self.cross_v, self.cross_out, + self.ffn1, self.ffn2, self.heads_hidden, self.heads_out, *self.norms): + m.release() + if self._tables is not None: + free_maps(*self._tables.values()) + self._tables = None + + def describe(self) -> Dict[str, Any]: + return {"name": self.name, "cells": self.cells, "classes": self.classes, "proposals": self.k, + "padded_k": self.kp, "dim": self.dim, "heads": [self.heads_self, self.heads_cross], + "head_dim": [self.hd_self, self.hd_cross], + "padded_head_dim": [padded_head_dim(self.hd_self), padded_head_dim(self.hd_cross)], + "scale": [self.scale_self, self.scale_cross], "attn_fp32_acc": self.attn_fp32_acc, + "head_slices": [list(s) for s in self.head_slices], "precision": self.policy.describe(), + "topk": self.select.describe()} + + +# ---------------------------------------------------------------------------------------------- oracle + +class QueryHeadOracle: + """fp32 (or fp64) torch oracle of the LITERAL head (no tables, no padding, no merged heads): the reference of + the device tests. Inputs NCHW ``lidar_feat (1, d, H, W)``, ``heat_nms (1, C, H, W)`` and the class-major + indices of the proposals.""" + + def __init__(self, weights: QueryHeadWeights, *, dtype: str = "float32"): + weights.validate() + self.w = weights + self.dtype = dtype + + def _t(self, a: Any): + import torch + + return torch.as_tensor(np.asarray(a, np.float64)).to(getattr(torch, self.dtype)) + + def _mha(self, a: MHAWeights, q_in, kv_in, v_in): + import torch + + t = self._t + q = q_in @ t(a.wq) + t(a.bq) + k = kv_in @ t(a.wk) + t(a.bk) + v = v_in @ t(a.wv) + t(a.bv) + h, d = int(a.num_heads), a.head_dim + q = q.reshape(-1, h, d).transpose(0, 1) + k = k.reshape(-1, h, d).transpose(0, 1) + v = v.reshape(-1, h, d).transpose(0, 1) + o = torch.softmax((q @ k.transpose(1, 2)) * float(a.scale), dim=-1) @ v + return o.transpose(0, 1).reshape(q_in.shape[0], h * d) @ t(a.wo) + t(a.bo) + + def _ln(self, i: int, x): + import torch + + g, b, eps = self.w.norms[i] + return torch.nn.functional.layer_norm(x, (x.shape[-1],), self._t(g), self._t(b), float(eps)) + + def _posembed(self, mlp: PosEmbedMLP, pos): + import torch + + t = self._t + return torch.relu(pos @ t(mlp.w0) + t(mlp.b0)) @ t(mlp.w1) + t(mlp.b1) + + def __call__(self, lidar_feat: Any, heat_nms: Any, indices: Any) -> Dict[str, np.ndarray]: + """-> ``{"query.feat" (K, d), "decoder.query_pos_embed", "decoder.self_attn", "decoder.cross_attn", + "decoder.out" (K, d), "heads" {name: (c, K)}, "query_heat" (C, K), "query_pos" (K, 2)}``.""" + import torch + + w = self.w + idx = np.asarray(indices).astype(np.int64).reshape(-1)[: w.num_proposals] + pos, cls = idx % w.cells, idx // w.cells + with torch.no_grad(): + feat = self._t(np.asarray(lidar_feat, np.float64).reshape(w.dim, w.cells).T) # (cells, d) + heat = np.asarray(heat_nms, np.float64).reshape(w.num_classes, w.cells) + x = feat[pos] + self._t(w.class_table)[cls] + qpos = self._t(np.asarray(w.bev_pos, np.float64)[pos]) + qpe = self._posembed(w.self_posembed, qpos) + key = feat + self._posembed(w.cross_posembed, self._t(w.bev_pos)) + s = x + qpe + x1 = self._ln(0, x + self._mha(w.self_attn, s, s, s)) + x2 = self._ln(1, x1 + self._mha(w.cross_attn, x1 + qpe, key, key)) + w1, b1, w2, b2 = (self._t(a) for a in w.ffn) + x3 = self._ln(2, x2 + torch.relu(x2 @ w1 + b1) @ w2 + b2) + heads = {} + for h in w.heads: + y = torch.relu(x3 @ self._t(h.w_hidden) + self._t(h.b_hidden)) @ self._t(h.w_out) + self._t(h.b_out) + heads[h.name] = y.T.numpy().astype(np.float32) + + def f(a): + return a.numpy().astype(np.float32) + + return {"query.feat": f(x), "decoder.query_pos_embed": f(qpe), "decoder.self_attn": f(x1), + "decoder.cross_attn": f(x2), "decoder.out": f(x3), "heads": heads, + "query_heat": heat[:, pos].astype(np.float32), "query_pos": f(qpos)} + + +# -------------------------------------------------------------------------------------------- assembly + +def assemble_transfusion(heads: Dict[str, np.ndarray], query_heat: Any, indices: Any, bev_pos: Any, *, + num_classes: int, num_proposals: int, cells: int) -> Dict[str, np.ndarray]: + """TransFusion's (mmdeploy-patched) graph outputs from the head's read-back arrays, in float32 on the host: + ``cls_score0 = query_heat * sigmoid(heads.heatmap)`` ``(1, C, K)``, ``bbox_pred0 = [center + query_pos, + height, dim, vel]`` ``(1, 8, K)``, ``dir_cls_pred0 = rot`` ``(1, 2, K)``; ``query_pos = bev_pos[pos]`` + (float32: never through bf16, S:transfusion:241). ``heads``: :meth:`TransFusionQueryHead.unpack_heads`; + ``query_heat``: the ``[.., Kp, 32]`` rows; ``indices``: the class-major proposal indices.""" + k = int(num_proposals) + idx = np.rint(np.asarray(indices, np.float64).reshape(-1)[:k]).astype(np.int64) + pos = idx % int(cells) + qpos = np.asarray(bev_pos, np.float32)[pos] # (K, 2) + qheat = np.asarray(query_heat, np.float32).reshape(-1, HEAD_PAD)[:k, :int(num_classes)].T + logits = np.asarray(heads["heatmap"], np.float32) + sig = (1.0 / (1.0 + np.exp(-logits.astype(np.float64)))).astype(np.float32) + cls = (qheat * sig).astype(np.float32)[None] + center = (np.asarray(heads["center"], np.float32) + qpos.T).astype(np.float32) + bbox = np.concatenate([center, heads["height"], heads["dim"], heads["vel"]], axis=0).astype(np.float32)[None] + return {"cls_score0": cls, "bbox_pred0": bbox, "dir_cls_pred0": np.asarray(heads["rot"], np.float32)[None]} diff --git a/code/tt_diffusion_planner/ttaw/nms.py b/code/tt_diffusion_planner/ttaw/nms.py new file mode 100644 index 0000000000000000000000000000000000000000..2cecbbdcd7971f95abcd11870f28f3654cd61308 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/nms.py @@ -0,0 +1,247 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C15 NMS family of the Autoware LiDAR detectors: circle NMS, rotated IoU-BEV NMS with perception_utils semantics, +and the area-based class remapper. + +- :func:`circle_nms` is ``circleNMS`` of ``autoware_lidar_centerpoint/lib/postprocess/circle_nms_kernel.cu:38-142``: + boxes in descending score order; a kept box suppresses every *later* box whose squared BEV centre distance is + strictly below ``threshold^2`` (float32: ``powf(dx, 2) + powf(dy, 2) < powf(thr, 2)``); greedy, class-agnostic. + A threshold <= 0 means 1.5 m in Autoware's config (``centerpoint_config.hpp:83-85,139``): the caller applies that. +- :func:`iou_bev_nms` is ``perception_utils::IouBevNms::apply`` (``perception_utils/src/iou_bev_nms.cpp:132-171``, + universe main 9ceaccf). Object ``t`` is compared with **every** earlier object ``s < t``, kept or suppressed + (S:centerpoint:291-295); pairs with different labels where one is PEDESTRIAN are skipped; pairs whose centre + distance squared exceeds ``search_distance_2d^2`` are skipped; the loop over ``s`` breaks at the first + ``iou > threshold``; ``t`` is kept iff the largest IoU seen is ``<= threshold``. ``sort=False`` (the CenterPoint + node's call) keeps the input order. IoU: the BEV rectangles of ``autoware_utils::toPolygon2d`` intersected + exactly (Sutherland-Hodgman on two convex quadrilaterals, float64), with the area guards 1e-6 (each box and the + intersection) and 0.01 (union). +- :class:`ClassRemapper` is ``perception_utils::DetectionClassRemapper::mapClasses`` + (``perception_utils/src/detection_class_remapper.cpp:54-70``): with ``bev_area = float(length * width)``, the first + destination label (in label order) whose matrix entry allows it and whose ``[min, max]`` contains the area wins. + +numpy only; no side effects on import. +""" +from __future__ import annotations + +import math +from dataclasses import dataclass +from typing import Any, List, Mapping, Optional, Sequence + +import numpy as np + +__all__ = [ + "PEDESTRIAN_LABEL", + "circle_nms", + "bev_box_corners", + "polygon_area", + "clip_convex", + "iou_bev", + "iou_bev_nms", + "ClassRemapper", +] + +PEDESTRIAN_LABEL = 7 # autoware_perception_msgs/ObjectClassification PEDESTRIAN + +_MIN_AREA = 1.0e-6 # iou_bev_nms.cpp compute2dIoU kMinArea +_MIN_UNION_AREA = 0.01 # kMinUnionArea + + +def circle_nms(x: Any, y: Any, dist_threshold: float) -> np.ndarray: + """Indices (ascending) of the boxes kept by Autoware's circle NMS. The input must already be in descending + score order (``thrust::sort``); see the module docstring.""" + xs = np.asarray(x, dtype=np.float32) + ys = np.asarray(y, dtype=np.float32) + n = len(xs) + thr2 = np.float32(dist_threshold) * np.float32(dist_threshold) + removed = np.zeros(n, dtype=bool) + keep: List[int] = [] + for i in range(n): + if removed[i]: + continue + keep.append(i) + dx = xs[i + 1:] - xs[i] + dy = ys[i + 1:] - ys[i] + removed[i + 1:] |= (dx * dx + dy * dy) < thr2 + return np.asarray(keep, dtype=np.int64) + + +def bev_box_corners(x: float, y: float, length: float, width: float, yaw: float) -> np.ndarray: + """(4, 2) float64 corners in ``toPolygon2d`` order: (+l/2, +w/2), (-l/2, +w/2), (-l/2, -w/2), (+l/2, -w/2) in + the object frame, rotated by ``yaw`` about the centre (counter-clockwise for positive sizes).""" + c, s = math.cos(yaw), math.sin(yaw) + hl, hw = length / 2.0, width / 2.0 + local = ((hl, hw), (-hl, hw), (-hl, -hw), (hl, -hw)) + return np.array([(x + c * px - s * py, y + s * px + c * py) for px, py in local], dtype=np.float64) + + +def polygon_area(poly: Any) -> float: + """Absolute shoelace area of a polygon (k, 2).""" + p = np.asarray(poly, dtype=np.float64) + if len(p) < 3: + return 0.0 + x, y = p[:, 0], p[:, 1] + return abs(float(np.dot(x, np.roll(y, -1)) - np.dot(np.roll(x, -1), y))) / 2.0 + + +def _signed_area(p: np.ndarray) -> float: + x, y = p[:, 0], p[:, 1] + return float(np.dot(x, np.roll(y, -1)) - np.dot(np.roll(x, -1), y)) / 2.0 + + +def clip_convex(subject: Any, clipper: Any) -> np.ndarray: + """Sutherland-Hodgman: ``subject`` polygon clipped by the convex polygon ``clipper`` (either orientation) -> + (k, 2) float64 intersection polygon (k may be 0).""" + clip = np.asarray(clipper, dtype=np.float64) + if _signed_area(clip) < 0: + clip = clip[::-1] + out = [tuple(p) for p in np.asarray(subject, dtype=np.float64)] + for i in range(len(clip)): + if not out: + break + ax, ay = clip[i] + bx, by = clip[(i + 1) % len(clip)] + inp, out = out, [] + + def inside(p): + return (bx - ax) * (p[1] - ay) - (by - ay) * (p[0] - ax) >= 0.0 + + def cross_point(p, q): + x1, y1, x2, y2 = p[0], p[1], q[0], q[1] + den = (x1 - x2) * (ay - by) - (y1 - y2) * (ax - bx) + if den == 0.0: + return q + t = ((x1 - ax) * (ay - by) - (y1 - ay) * (ax - bx)) / den + return (x1 + t * (x2 - x1), y1 + t * (y2 - y1)) + + prev = inp[-1] + for cur in inp: + if inside(cur): + if not inside(prev): + out.append(cross_point(prev, cur)) + out.append(cur) + elif inside(prev): + out.append(cross_point(prev, cur)) + prev = cur + return np.asarray(out, dtype=np.float64).reshape(-1, 2) + + +def _iou_from(pa: np.ndarray, area_a: float, env_a: np.ndarray, pb: np.ndarray, area_b: float, + env_b: np.ndarray) -> float: + if area_a < _MIN_AREA or area_b < _MIN_AREA: + return 0.0 + if env_a[0] > env_b[2] or env_b[0] > env_a[2] or env_a[1] > env_b[3] or env_b[1] > env_a[3]: + return 0.0 # disjoint envelopes (a speed guard in perception_utils; the intersection would be empty) + inter = polygon_area(clip_convex(pa, pb)) + if inter < _MIN_AREA: + return 0.0 + union = area_a + area_b - inter + if union < _MIN_UNION_AREA: + return 0.0 + return min(1.0, inter / union) + + +def _envelope(p: np.ndarray) -> np.ndarray: + return np.array([p[:, 0].min(), p[:, 1].min(), p[:, 0].max(), p[:, 1].max()]) + + +def iou_bev(a: Sequence[float], b: Sequence[float]) -> float: + """Rotated BEV IoU of two boxes ``(x, y, length, width, yaw)`` with perception_utils' ``compute2dIoU`` guards.""" + pa, pb = bev_box_corners(*a), bev_box_corners(*b) + return _iou_from(pa, polygon_area(pa), _envelope(pa), pb, polygon_area(pb), _envelope(pb)) + + +def iou_bev_nms(x: Any, y: Any, length: Any, width: Any, yaw: Any, labels: Any, *, search_distance_2d: float = 10.0, + iou_threshold: float = 0.1, scores: Optional[Any] = None, sort: bool = False, + pedestrian_label: int = PEDESTRIAN_LABEL) -> np.ndarray: + """Indices of the objects kept by ``IouBevNms::apply`` (module docstring), in processing order. ``labels`` are + ObjectClassification labels; ``yaw`` is the ROS yaw. ``sort=True`` first orders by descending ``scores`` + (``std::stable_sort``).""" + if not math.isfinite(search_distance_2d) or search_distance_2d < 0.0: + raise ValueError("search_distance_2d must be a finite non-negative value") + if not math.isfinite(iou_threshold) or not 0.0 <= iou_threshold <= 1.0: + raise ValueError("iou_threshold must be a finite value between 0 and 1") + xs = np.asarray(x, dtype=np.float64) + n = len(xs) + order = np.arange(n) + if sort: + if scores is None: + raise ValueError("sort=True needs scores") + order = np.argsort(-np.asarray(scores, dtype=np.float64), kind="stable") + ys = np.asarray(y, dtype=np.float64)[order] + xs = xs[order] + ls = np.asarray(length, dtype=np.float64)[order] + ws = np.asarray(width, dtype=np.float64)[order] + yw = np.asarray(yaw, dtype=np.float64)[order] + lab = np.asarray(labels, dtype=np.int64)[order] + polys = [bev_box_corners(xs[i], ys[i], ls[i], ws[i], yw[i]) for i in range(n)] + areas = [polygon_area(p) for p in polys] + envs = [_envelope(p) for p in polys] + sd2 = float(search_distance_2d) * float(search_distance_2d) + keep: List[int] = [] + for t in range(n): + max_iou = 0.0 + for s in range(t): + if lab[t] != lab[s] and (lab[t] == pedestrian_label or lab[s] == pedestrian_label): + continue + dx, dy = xs[t] - xs[s], ys[t] - ys[s] + if not dx * dx + dy * dy <= sd2: + continue + iou = _iou_from(polys[t], areas[t], envs[t], polys[s], areas[s], envs[s]) + max_iou = max(max_iou, iou) + if iou > iou_threshold: + break + if max_iou <= iou_threshold: + keep.append(int(order[t])) + return np.asarray(keep, dtype=np.int64) + + +@dataclass(frozen=True) +class ClassRemapper: + """``DetectionClassRemapper`` with square ``allow`` (bool), ``min_area`` and ``max_area`` matrices indexed + ``[source_label, destination_label]`` over the ObjectClassification labels. Non-finite bounds become + ``DBL_MAX`` (``detection_class_remapper.cpp:43-48``).""" + + allow: np.ndarray + min_area: np.ndarray + max_area: np.ndarray + + @classmethod + def from_lists(cls, allow_remapping_by_area_matrix: Sequence[int], min_area_matrix: Sequence[float], + max_area_matrix: Sequence[float]) -> "ClassRemapper": + """The three flat lists of ``detection_class_remapper.param.yaml`` (row = source, column = destination).""" + if not (len(allow_remapping_by_area_matrix) == len(min_area_matrix) == len(max_area_matrix)): + raise ValueError("Class remapping matrices must have equal sizes.") + n = int(math.isqrt(len(min_area_matrix))) + if n == 0 or n * n != len(min_area_matrix): + raise ValueError("Class remapping matrices must be non-empty and square.") + + def finite(v: Sequence[float]) -> np.ndarray: + a = np.asarray(v, dtype=np.float64) + return np.where(np.isfinite(a), a, np.finfo(np.float64).max).reshape(n, n) + + return cls(np.asarray(allow_remapping_by_area_matrix, dtype=np.int64).reshape(n, n) != 0, + finite(min_area_matrix), finite(max_area_matrix)) + + @classmethod + def from_params(cls, params: Mapping[str, Any]) -> "ClassRemapper": + """From the ``ros__parameters`` mapping of ``detection_class_remapper.param.yaml``.""" + return cls.from_lists(params["allow_remapping_by_area_matrix"], params["min_area_matrix"], + params["max_area_matrix"]) + + @property + def num_labels(self) -> int: + return int(self.allow.shape[0]) + + def apply(self, labels: Any, length: Any, width: Any) -> np.ndarray: + """New labels (uint8) for objects with ``labels`` and BEV ``length`` x ``width``.""" + lab = np.asarray(labels, dtype=np.int64).copy() + area = (np.asarray(length, dtype=np.float64) * np.asarray(width, dtype=np.float64)).astype(np.float32) + area = area.astype(np.float64) + for i in range(len(lab)): + src = int(lab[i]) + if src < 0 or src >= self.num_labels: + continue + for dst in range(self.num_labels): + if self.allow[src, dst] and self.min_area[src, dst] <= area[i] <= self.max_area[src, dst]: + lab[i] = dst + break + return lab.astype(np.uint8) diff --git a/code/tt_diffusion_planner/ttaw/ops/__init__.py b/code/tt_diffusion_planner/ttaw/ops/__init__.py new file mode 100644 index 0000000000000000000000000000000000000000..0e5e3c82a25bd65d8a03b4e4354929c839d9e545 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/ops/__init__.py @@ -0,0 +1,28 @@ +# SPDX-License-Identifier: Apache-2.0 +"""Custom-op helpers. Kernel ``.cpp`` / ``.hpp`` files are package data under ``ops/kernels/`` (shipped through +``source.extra_code`` and ``[tool.setuptools.package-data]``), and are always referenced by absolute path, because +tt-metal resolves a relative kernel path against the CWD, ``TT_METAL_KERNEL_PATH`` and ``TT_METAL_HOME`` first +(TT_PLATFORM.md section 5.1). The C17-C23 op modules and the K1-K11 kernels of PLAN.md section 1.3 are added here +by the first port that needs them, with a numpy oracle test and the hang protocol of PLAN.md section 4.4. + +Op modules (import them explicitly; this package imports none of them): ``conv`` (C17 conv builders), ``upsample`` +(C18 up-sampling and interpolation), ``attention`` (C20 SDPA wrapper), ``gather`` (C19 row gathers), ``topk`` (C21 +top-k selection), ``heatmap`` (C22 heatmap local max), ``deform`` (C23 grid_sample helpers), ``segment`` (K1 +``segment_reduce``: the first ``generic_op`` kernel, ``kernels/segment_reduce_dm.cpp``). +""" +from __future__ import annotations + +import os + +__all__ = ["KERNEL_DIR", "kernel_path"] + +KERNEL_DIR = os.path.join(os.path.dirname(os.path.abspath(__file__)), "kernels") + + +def kernel_path(name: str, base: str = KERNEL_DIR) -> str: + """Absolute path of a kernel source shipped with the package (``kernel_path("segment_max_reader.cpp")``). + Raises ``FileNotFoundError`` early instead of tt-metal's "doesn't exist in any of the searched paths".""" + path = os.path.join(base, name) + if not os.path.isfile(path): + raise FileNotFoundError(f"kernel source {path} is missing (not shipped as package data?)") + return path diff --git a/code/tt_diffusion_planner/ttaw/ops/attention.py b/code/tt_diffusion_planner/ttaw/ops/attention.py new file mode 100644 index 0000000000000000000000000000000000000000..7776d9ba671a7afd7ccc787152e7bb95e43dc9ec --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/ops/attention.py @@ -0,0 +1,386 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C20: attention on the p150 through ``ttnn.transformer.scaled_dot_product_attention`` (PLAN.md 1.2 C20, probe P7). + +SDPA has three silent traps (common/PROBES.md P7; ``probes/g4-attention-gridsample-msda.md``): + +- ``is_causal`` defaults to True: with Sq == Sk it silently applies a causal mask, with Sq != Sk it is rejected; +- the default ``scale`` is ``1/sqrt(padded head_dim)``: after a 16 -> 32 head padding it is wrong (PCC 0.979); +- the default program config (32/32 chunks) is 2.5-7x slower everywhere and inaccurate over >= 32k keys. + +:func:`sdpa` is the only way the ports call the op: ``is_causal=False`` always, ``scale`` is a required keyword, +and a ``program_config`` is always passed (chunk sizes from :func:`chunk_config`, the P7 table). Without a mask, +pass the logical, unaligned Sq / Sk: the op pads to its chunks and masks its own padding element by element. A user +mask is an additive bf16 TILE tensor **in DRAM**, ``[1|B, 1|H, Sq, Sk]``; ``-inf`` is safe, also when a whole first +key chunk is masked (P7). + +**A user mask needs a tile-aligned Sk** (found by the Diffusion Planner port, 2026-10-07): with a mask, the reader +fills only the key *tiles* past ``ceil(Sk / 32)`` with ``-inf`` and reads the last partial tile from the mask +itself, whose tile padding is 0 after a host tilize (``reader_interleaved.cpp``, "Mask read"). The up to 31 padded +keys then join the softmax with score ``q . k_pad`` and value ``v_pad``: rel-L2 0.18 instead of 0.025 at 321 x 321 +with 89 valid keys, 0.06 at P7's masked 564 x 564. :func:`sdpa` therefore refuses a mask with ``Sk % 32 != 0``: +give K / V a tile-aligned key count (:func:`aligned_keys`; the extra keys are masked with ``-inf`` in the mask, +``key_bias(padded_validity, sq)``), which costs nothing because the tiles are padded anyway. + +Head layout. The op wants ``[B, H, S, D]`` bf16 TILE interleaved tensors with ``D`` a multiple of 32 and no padding on +B / H / D. :func:`split_heads` / :func:`split_qkv` / :func:`split_q_kv` turn projection outputs ``[B, 1, S, H*D]`` +into heads (``ttnn.experimental.nlp_create_qkv_heads``, one program, logical S kept), and ``sdpa(..., +concat_heads=True)`` (the op's ``output_concat_heads``) or :func:`merge_heads` go back to ``[B, 1, S, H*D]``. +A head dim below 32 (16 in TransFusion / BEVFusion / PTv3) is zero-padded *in the weights* +(:func:`pad_head_columns` on the Q / K / V projections, :func:`pad_head_rows` on the output projection): exact +(padded output columns are 0, P7) and free at run time; pass ``scale=1/sqrt(logical head dim)``. + +Masks. :func:`key_bias` builds ``[1, 1, Sq, Sk]`` (0 for valid keys, ``-inf`` otherwise) on the host; inside a +trace :func:`expand_key_bias` repeats a persistent ``[1, 1, 1, Sk]`` bias row (:func:`key_bias_row`) to +``[1, 1, Sq, Sk]`` in DRAM (one program, the upload shrinks from Sq * Sk to Sk values). + +Fallback. :func:`attention_matmul` is plain ``softmax(scale Q K^T + bias) V`` with matmuls (for dtypes or layouts the +op rejects, e.g. fp32 attention); it needs a tile-aligned Sk (pad the keys and mask them with ``-inf``). +:func:`attention_reference` is the float64 numpy oracle of both. + +Importing this module never imports ttnn (ttaw rule); ttnn is imported inside the functions that need it. +""" +from __future__ import annotations + +import math +from typing import Any, Dict, Optional, Sequence, Tuple + +import numpy as np + +__all__ = [ + "aligned_keys", + "HEAD_DIM_MULTIPLE", + "PROGRAM_CONFIG_MIN_SK", + "CHUNK_TABLE", + "padded_head_dim", + "chunk_config", + "sdpa_program_config", + "sdpa_compute_config", + "sdpa", + "split_heads", + "split_qkv", + "split_q_kv", + "merge_heads", + "key_bias", + "key_bias_row", + "expand_key_bias", + "pad_head_dim", + "pad_head_columns", + "pad_head_rows", + "attention_reference", + "attention_matmul", +] + +HEAD_DIM_MULTIPLE = 32 # SDPA rejects a padded head dim ("Padding is not supported on the head_dim dimension") +TILE = 32 +# P7: from about 1,000 keys on, the default 32/32 chunks are 2.5-7x slower; this wrapper passes a config always. +PROGRAM_CONFIG_MIN_SK = 1000 + +# Exact (Sq, Sk) shapes measured on the p150b (ETH, 12x10): (q_chunk, k_chunk) and the evidence. Shapes not listed +# use the Sk-range rules of :func:`chunk_config`. Ports add their measured shapes here (with the log). +CHUNK_TABLE: Dict[Tuple[int, int], Tuple[int, int, str]] = { + (500, 500): (128, 128, "P7 tf_self 0.029 ms (default 0.047)"), + (512, 512): (512, 512, "P7 pt_patch_s0/s4 0.172-0.222 ms (q128 k128 0.413-0.571)"), + (564, 564): (64, 64, "P7 dp_fusion_masked 0.058 ms (default 0.094, q128 k128 0.158)"), + (321, 564): (64, 64, "P7 dp_cross 0.029 ms (default 0.030)"), + (576, 576): (64, 64, "DP fusion, masked: 0.061-0.092 ms (q32 k64 0.117, q32 k32 0.129, q128 k128 0.275)"), + (352, 352): (64, 64, "DP DiT self, masked: 0.063 ms (q64 k128 0.059, q32 k32 0.080, q128 k128 0.071)"), + (352, 564): (64, 64, "DP DiT cross: 0.033-0.043 ms (q32 k32 0.035, q128 k128 0.052, q64 k128 0.077)"), + (900, 1668): (64, 128, "P7 sp_self 0.052 ms (default 0.139)"), + (900, 6000): (64, 512, "P7 sp_cross 0.106 ms (default 0.469)"), +} + + +def padded_head_dim(head_dim: int, multiple: int = HEAD_DIM_MULTIPLE) -> int: + """The head dim SDPA sees: ``head_dim`` rounded up to a multiple of 32.""" + if head_dim <= 0: + raise ValueError(f"head_dim must be positive, got {head_dim}") + return -(-int(head_dim) // multiple) * multiple + + +def aligned_keys(n: int, multiple: int = TILE) -> int: + """The tile-aligned key count a masked attention over ``n`` real keys runs on (the extra keys are masked).""" + if n <= 0: + raise ValueError(f"key count must be positive, got {n}") + return -(-int(n) // multiple) * multiple + + +def chunk_config(sq: int, sk: int) -> Tuple[int, int]: + """``(q_chunk, k_chunk)`` for logical ``Sq`` x ``Sk`` (multiples of 32, as the op requires). + + Exact shapes come from :data:`CHUNK_TABLE`. Otherwise, by key length (P7: chunk sizes are what matters for + long keys; q64 keeps ``B * H * Sq / 64`` work units near the 120 cores at Sq ~ 500-900): + Sk >= 16,384 -> q64 k1024 (P7 tf / bf / pt cross 0.45-0.88 ms; q64 k512 + ``fp32_acc`` halves the max error at + ~1.4x the time); Sk >= 4,096 -> q64 k512; Sk >= 1,000 -> q64 k128; shorter keys -> q64 k64 (32 when the + sequence fits one tile). + """ + sq, sk = int(sq), int(sk) + if sq <= 0 or sk <= 0: + raise ValueError(f"sequence lengths must be positive, got Sq={sq}, Sk={sk}") + if (sq, sk) in CHUNK_TABLE: + q, k, _ = CHUNK_TABLE[(sq, sk)] + return q, k + q = 32 if sq <= TILE else 64 + if sk >= 16384: + return q, 1024 + if sk >= 4096: + return q, 512 + if sk >= PROGRAM_CONFIG_MIN_SK: + return q, 128 + return q, (32 if sk <= TILE else 64) + + +def sdpa_program_config(device: Any, sq: int, sk: int, *, q_chunk: Optional[int] = None, k_chunk: Optional[int] = None, + exp_approx_mode: Optional[bool] = None): + """``ttnn.SDPAProgramConfig`` over the device's whole compute grid (read from the device, never hard-coded) + with the :func:`chunk_config` chunks unless ``q_chunk`` / ``k_chunk`` are given.""" + import ttnn + + q_default, k_default = chunk_config(sq, sk) + q_chunk = int(q_chunk or q_default) + k_chunk = int(k_chunk or k_default) + for name, value in (("q_chunk", q_chunk), ("k_chunk", k_chunk)): + if value <= 0 or value % TILE: + raise ValueError(f"{name}={value}: must be a positive multiple of {TILE}") + return ttnn.SDPAProgramConfig(compute_with_storage_grid_size=device.compute_with_storage_grid_size(), + q_chunk_size=q_chunk, k_chunk_size=k_chunk, exp_approx_mode=exp_approx_mode) + + +def sdpa_compute_config(*, fp32_acc: bool = False, fidelity: str = "HiFi2", approx: bool = True): + """Explicit SDPA compute config. The default equals the op's own (HiFi2, bf16 DST, approximate exp): P7 found + neither HiFi4 nor exact exp more accurate; ``fp32_acc=True`` (the op's standard instead of its streaming kernel) + halves the max error (0.11 -> 0.056 over 36k keys) at 1.35-1.45x the time.""" + from ..precision import compute_kernel_config + + return compute_kernel_config(fidelity, fp32_acc=fp32_acc, approx=approx) + + +def _seq(t: Any) -> int: + return int(t.shape[-2]) + + +def sdpa(q: Any, k: Any, v: Any, *, scale: float, attn_mask: Any = None, program_config: Any = "auto", + compute_kernel_config: Any = None, fp32_acc: bool = False, concat_heads: bool = False, + memory_config: Any = None): + """Non-causal scaled dot-product attention ``softmax(scale * Q K^T + attn_mask) V``. + + ``q`` ``[B, H, Sq, D]``, ``k`` / ``v`` ``[B, Hkv, Sk, D]`` (bf16 / bfp8 / bfp4, TILE, interleaved, D a multiple of + 32); ``scale`` is required: ``1/sqrt(logical head dim)`` (never the op's padded-head default). ``attn_mask``: + additive ``[1|B, 1|H, Sq, Sk]`` bf16 TILE tensor in DRAM, or None; with a mask ``Sk`` must be a multiple of 32 + (module docstring). ``program_config``: ``"auto"`` (the + :func:`chunk_config` chunks over the whole grid), an ``SDPAProgramConfig``, or None (op default; refused once + Sk >= :data:`PROGRAM_CONFIG_MIN_SK`, P7). ``compute_kernel_config`` overrides :func:`sdpa_compute_config` + (``fp32_acc``). ``concat_heads=True`` returns ``[B, 1, Sq, H*D]`` instead of ``[B, H, Sq, D]`` (no separate + concat program). Returns the device tensor. + """ + import ttnn + + scale = float(scale) + if not math.isfinite(scale) or scale <= 0.0: + raise ValueError(f"scale must be a positive finite float, got {scale!r}") + if len(q.shape) != 4 or len(k.shape) != 4 or len(v.shape) != 4: + raise ValueError(f"sdpa expects [B, H, S, D] tensors, got q{tuple(q.shape)} k{tuple(k.shape)} " + f"v{tuple(v.shape)}") + d = int(q.shape[-1]) + if d % HEAD_DIM_MULTIPLE: + raise ValueError(f"head dim {d} is not a multiple of {HEAD_DIM_MULTIPLE}: zero-pad it in the projection " + "weights (pad_head_columns / pad_head_rows) and pass scale=1/sqrt(logical head dim)") + sq, sk = _seq(q), _seq(k) + if attn_mask is not None and sk % TILE: + raise ValueError(f"a user mask needs a tile-aligned key count (Sk={sk}): the op masks the padded keys of the " + "last partial tile only through the mask's own tile padding (0 after a host tilize), so " + f"they would join the softmax. Run on aligned_keys({sk}) = {aligned_keys(sk)} keys and mask " + "the extra ones with -inf (key_bias on the padded validity)") + if isinstance(program_config, str): + if program_config != "auto": + raise ValueError(f"program_config={program_config!r}: expected 'auto', a config object or None") + program_config = sdpa_program_config(q.device(), sq, sk) + elif program_config is None and sk >= PROGRAM_CONFIG_MIN_SK: + raise ValueError(f"Sk={sk} >= {PROGRAM_CONFIG_MIN_SK}: the op's default 32/32 chunks are slow and inaccurate " + "for long keys (probe P7); pass program_config='auto' or an SDPAProgramConfig") + if compute_kernel_config is None: + compute_kernel_config = sdpa_compute_config(fp32_acc=fp32_acc) + kwargs: Dict[str, Any] = {} + if concat_heads: + kwargs["output_concat_heads"] = True + return ttnn.transformer.scaled_dot_product_attention( + q, k, v, attn_mask=attn_mask, is_causal=False, scale=scale, memory_config=memory_config, + program_config=program_config, compute_kernel_config=compute_kernel_config, **kwargs) + + +# ---------------------------------------------------------------------------------------------------- heads + +def _check_4d_heads(x: Any, what: str) -> None: + if len(x.shape) != 4 or int(x.shape[1]) != 1: + raise ValueError(f"{what} must be [B, 1, S, width] (a 4-D projection output), got {tuple(x.shape)}") + + +def split_heads(x: Any, num_heads: int, *, memory_config: Any = None): + """``[B, 1, S, H*D]`` -> ``[B, H, S, D]`` (one ``nlp_create_qkv_heads`` program with no K / V heads; the logical + S is kept, so unaligned sequences stay unaligned for :func:`sdpa`).""" + import ttnn + + _check_4d_heads(x, "split_heads input") + width = int(x.shape[-1]) + if width % num_heads or (width // num_heads) % HEAD_DIM_MULTIPLE: + raise ValueError(f"width {width} does not split into {num_heads} heads of a multiple of {HEAD_DIM_MULTIPLE}") + out, _, _ = ttnn.experimental.nlp_create_qkv_heads(x, num_heads=num_heads, num_kv_heads=0, + transpose_k_heads=False, memory_config=memory_config) + return out + + +def split_qkv(qkv: Any, num_heads: int, *, num_kv_heads: Optional[int] = None, memory_config: Any = None): + """Fused ``[B, 1, S, (H + 2 Hkv) * D]`` (Q | K | V column blocks, heads contiguous) -> ``(q, k, v)`` + ``[B, H|Hkv, S, D]`` in one program (K not transposed).""" + import ttnn + + _check_4d_heads(qkv, "split_qkv input") + kv_heads = num_heads if num_kv_heads is None else int(num_kv_heads) + width = int(qkv.shape[-1]) + if width % (num_heads + 2 * kv_heads) or (width // (num_heads + 2 * kv_heads)) % HEAD_DIM_MULTIPLE: + raise ValueError(f"width {width} does not split into {num_heads} + 2 x {kv_heads} heads of a multiple of " + f"{HEAD_DIM_MULTIPLE}") + return ttnn.experimental.nlp_create_qkv_heads(qkv, num_heads=num_heads, num_kv_heads=kv_heads, + transpose_k_heads=False, memory_config=memory_config) + + +def split_q_kv(q: Any, kv: Any, num_heads: int, *, num_kv_heads: Optional[int] = None, memory_config: Any = None): + """Separate ``q`` ``[B, 1, S, H*D]`` and ``kv`` ``[B, 1, S, 2*Hkv*D]`` (K | V column blocks) with the SAME S + (self-attention with Q and K / V from different inputs, e.g. a pre-LN Q next to an un-normalised K / V) + -> ``(q, k, v)`` heads in one program. Cross-attention with another key length: :func:`split_heads` each.""" + import ttnn + + _check_4d_heads(q, "split_q_kv q") + _check_4d_heads(kv, "split_q_kv kv") + if int(q.shape[-2]) != int(kv.shape[-2]): + raise ValueError(f"split_q_kv needs the same sequence length (Q {q.shape[-2]}, K/V {kv.shape[-2]})") + kv_heads = num_heads if num_kv_heads is None else int(num_kv_heads) + return ttnn.experimental.nlp_create_qkv_heads(q, kv, num_heads=num_heads, num_kv_heads=kv_heads, + transpose_k_heads=False, memory_config=memory_config) + + +def merge_heads(x: Any, *, memory_config: Any = None): + """``[B, H, S, D]`` -> ``[B, 1, S, H*D]`` (``nlp_concat_heads``). Prefer ``sdpa(..., concat_heads=True)``.""" + import ttnn + + if len(x.shape) != 4: + raise ValueError(f"merge_heads expects [B, H, S, D], got {tuple(x.shape)}") + return ttnn.experimental.nlp_concat_heads(x, memory_config=memory_config) + + +# ---------------------------------------------------------------------------------------------------- masks + +def key_bias_row(valid: Sequence[bool], *, dtype: Any = np.float32) -> np.ndarray: + """``[1, 1, 1, Sk]`` additive key bias: 0 for valid keys, ``-inf`` for masked ones (bf16-exact values).""" + v = np.asarray(valid, dtype=bool).reshape(-1) + if v.size == 0: + raise ValueError("key_bias_row: empty validity vector") + if not v.any(): + raise ValueError("key_bias_row: every key is masked (the softmax of a fully masked row is undefined)") + return np.where(v, 0.0, -np.inf).astype(dtype).reshape(1, 1, 1, -1) + + +def key_bias(valid: Sequence[bool], sq: int, *, dtype: Any = np.float32) -> np.ndarray: + """``[1, 1, Sq, Sk]`` additive key mask for :func:`sdpa` (identical rows), built on the host. Upload it as a + bf16 TILE tensor in DRAM (the op's requirement), or keep a :func:`key_bias_row` input and expand it on the + device with :func:`expand_key_bias`.""" + row = key_bias_row(valid, dtype=dtype) + return np.ascontiguousarray(np.broadcast_to(row, (1, 1, int(sq), row.shape[-1]))) + + +def expand_key_bias(row: Any, sq: int, *, memory_config: Any = None): + """Device-side :func:`key_bias`: repeat a ``[1, 1, 1, Sk]`` bf16 TILE bias row to ``[1, 1, Sq, Sk]`` in DRAM + (one program; traceable).""" + import ttnn + + if len(row.shape) != 4 or int(row.shape[-2]) != 1: + raise ValueError(f"expand_key_bias expects a [1, 1, 1, Sk] row, got {tuple(row.shape)}") + return ttnn.repeat(row, [1, 1, int(sq), 1], + memory_config=ttnn.DRAM_MEMORY_CONFIG if memory_config is None else memory_config) + + +# ---------------------------------------------------------------------------------------------------- head padding + +def pad_head_dim(x: Any, padded: Optional[int] = None) -> np.ndarray: + """Zero-pad the last (head) dim of a host array ``[..., D]`` to ``padded`` (default :func:`padded_head_dim`).""" + a = np.asarray(x) + d = a.shape[-1] + p = padded_head_dim(d) if padded is None else int(padded) + if p < d: + raise ValueError(f"padded head dim {p} < head dim {d}") + if p == d: + return a + return np.pad(a, [(0, 0)] * (a.ndim - 1) + [(0, p - d)]) + + +def pad_head_columns(w: Any, b: Any, num_heads: int, head_dim: int, + padded: Optional[int] = None) -> Tuple[np.ndarray, Optional[np.ndarray]]: + """Projection ``[in, H*D]`` (``y = x @ w + b``) -> ``[in, H*Dp]`` with zero columns after each head's D (and the + bias likewise), so the projection itself emits padded heads (exact: the padded columns are 0).""" + w = np.asarray(w) + p = padded_head_dim(head_dim) if padded is None else int(padded) + if w.shape[-1] != num_heads * head_dim: + raise ValueError(f"weight has {w.shape[-1]} columns, expected {num_heads} x {head_dim}") + wh = w.reshape(w.shape[:-1] + (num_heads, head_dim)) + wp = np.zeros(w.shape[:-1] + (num_heads, p), w.dtype) + wp[..., :head_dim] = wh + bp = None + if b is not None: + b = np.asarray(b).reshape(num_heads, head_dim) + bp = np.zeros((num_heads, p), b.dtype) + bp[:, :head_dim] = b + bp = bp.reshape(-1) + return wp.reshape(w.shape[:-1] + (num_heads * p,)), bp + + +def pad_head_rows(w: Any, num_heads: int, head_dim: int, padded: Optional[int] = None) -> np.ndarray: + """Output projection ``[H*D, out]`` -> ``[H*Dp, out]`` with zero rows at each head's padded positions, matching + the ``[.., H*Dp]`` concatenated-heads output of a padded attention.""" + w = np.asarray(w) + p = padded_head_dim(head_dim) if padded is None else int(padded) + if w.shape[0] != num_heads * head_dim: + raise ValueError(f"weight has {w.shape[0]} rows, expected {num_heads} x {head_dim}") + wp = np.zeros((num_heads, p) + w.shape[1:], w.dtype) + wp[:, :head_dim] = w.reshape((num_heads, head_dim) + w.shape[1:]) + return wp.reshape((num_heads * p,) + w.shape[1:]) + + +# ---------------------------------------------------------------------------------------------------- oracle / fallback + +def attention_reference(q: Any, k: Any, v: Any, *, scale: float, bias: Any = None) -> np.ndarray: + """float64 numpy ``softmax(scale Q K^T + bias) V`` over ``[B, H, S, D]`` arrays (K / V heads broadcast when + ``Hkv`` divides ``H``); ``bias`` broadcasts to ``[B, H, Sq, Sk]`` (``-inf`` allowed).""" + q = np.asarray(q, np.float64) + k = np.asarray(k, np.float64) + v = np.asarray(v, np.float64) + if k.shape[1] != q.shape[1]: + rep = q.shape[1] // k.shape[1] + k, v = np.repeat(k, rep, axis=1), np.repeat(v, rep, axis=1) + s = np.einsum("bhqd,bhkd->bhqk", q, k) * float(scale) + if bias is not None: + s = s + np.asarray(bias, np.float64) + s = s - s.max(axis=-1, keepdims=True) + p = np.exp(s) + p /= p.sum(axis=-1, keepdims=True) + return np.einsum("bhqk,bhkd->bhqd", p, v) + + +def attention_matmul(q: Any, k: Any, v: Any, *, scale: float, attn_mask: Any = None, + compute_kernel_config: Any = None): + """Small-MHA fallback for what :func:`sdpa` rejects (e.g. fp32 Q / K / V): ``softmax(scale Q K^T + mask) V`` + with two matmuls. ``q`` ``[B, H, Sq, D]``, ``k`` / ``v`` ``[B, H, Sk, D]`` with a tile-aligned ``Sk`` (pad the + keys and mask the padding with ``-inf``: a softmax over tile padding would include it). Materialises the + ``[B, H, Sq, Sk]`` scores, so keep it for small shapes.""" + import ttnn + + from ..precision import compute_kernel_config as ckc + + sk = int(k.shape[-2]) + if sk % TILE: + raise ValueError(f"attention_matmul needs a tile-aligned Sk (got {sk}): pad the keys and mask them") + cfg = compute_kernel_config if compute_kernel_config is not None else ckc("HiFi4", fp32_acc=True) + scores = ttnn.matmul(q, k, transpose_b=True, compute_kernel_config=cfg) + scores = ttnn.multiply(scores, float(scale)) + if attn_mask is not None: + scores = ttnn.add(scores, attn_mask) + probs = ttnn.softmax(scores, dim=-1, numeric_stable=True, compute_kernel_config=cfg) + return ttnn.matmul(probs, v, compute_kernel_config=cfg) diff --git a/code/tt_diffusion_planner/ttaw/ops/conv.py b/code/tt_diffusion_planner/ttaw/ops/conv.py new file mode 100644 index 0000000000000000000000000000000000000000..4ee4ffd6c87b24912b49ce7006353cd787869932 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/ops/conv.py @@ -0,0 +1,771 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C17: conv2d builders for NHWC feature maps (PLAN.md section 1.2 row C17; probes P5 and P14, common/PROBES.md). + +A conv layer of a port is built once from host weights and called on device feature maps, eagerly for the warm-up +and then inside trace captures. What every builder here guarantees: + +- **Weights prepared once.** The first call hands the host weights to ``ttnn.conv2d`` (which prepares them for the + chosen sharding and moves them to the device, a host write); the prepared device tensors are kept per input spec + and reused, so a traced call never writes from the host ("Writes are not supported during trace capture", + RP section 1.4). Call each layer once eagerly before capturing (``TraceRunner`` warm-up does). +- **Explicit precision.** ``compute_config`` comes from :mod:`..precision` (default :data:`CONV_PRECISION`: HiFi4 + + fp32 accumulation + packer L1 accumulation, exact SFPU activations; never ttnn's implicit defaults; without the L1 + accumulation the partial sums of wide convs round to bf16 between inner blocks). Weights default to bf16 + (``Precision.weights`` selects bfp8); ``output_dtype`` may be float32 (e.g. heads whose logits feed thresholds). +- **Fused bias + activation** in the conv epilogue: ``relu`` (packer), ``relu6`` / ``silu`` / ``gelu`` / + ``sigmoid`` (SFPU). RELU6 is verified on the device (P5: output in [0, 6], PCC >= 0.99997). HARDSWISH is refused: + the fused form is silently skipped on this tree (P5; apply ``ttnn.hardswish`` as a separate op). +- **Grid from the device.** Nothing here hard-codes 12x10 / 11x10: ttnn's auto-sharding reads + ``device.compute_with_storage_grid_size()``; explicit shard overrides go through ``conv_config``. +- **Memory placement.** ``slicing="auto"`` keeps ttnn's rule: a DRAM input runs the DRAM-sliced path with automatic + slice counts (P14: every accepted config captures and replays bit-exactly; explicit counts are often rejected, so + never hard-code them) and writes a DRAM-interleaved output; an L1 input runs the L1 path. 1x1 stride-1 convs are + matmuls and always run the L1 path; ``output_memory_config`` (default ``"dram"``) brings their output back to DRAM + so a graph can keep every activation in DRAM between ops (the robust layout of a first port; L1 residency is an + optimization). Feed large-kernel convs TILE tensors (``[1, 1, N*H*W, C]``, what a producing conv writes): from + row-major DRAM they cost 2x (P5: 926 vs 482 us for YOLOX's 12 large-kernel convs). + +Feature maps travel as :class:`FeatureMap` (device tensor + logical N, H, W, C), because ttnn's flattened +``[1, 1, N*H*W, C]`` layout no longer carries the spatial size. + +Builders: :class:`Conv2d` (any kernel / stride / padding / dilation / groups), :class:`KSplitConv` +(``conv(concat(x_1..x_n))`` as ``sum_i conv_i(x_i)`` without building the concat; P14: 1.7-4.1x faster and more +accurate for CenterPoint / PointPainting / FRNet's wide concats), :class:`ConvTranspose2d` for kernel == stride +(P5: k = 2 -> ``ttnn.conv_transpose2d``; k >= 3 -> ``ttnn.linear`` + depth-to-space, 5x faster at k = 4), and the +host helpers :func:`merge_sibling_convs` (one conv for sibling heads that read the same input) / +:func:`split_k_weights`. Glue: :func:`concat_channels`, :func:`split_channels`, :func:`residual_add`, +:func:`apply_activation`, :func:`fmap_from_numpy` / :func:`fmap_to_numpy` (NCHW <-> device NHWC). +""" +from __future__ import annotations + +from dataclasses import dataclass +from typing import Any, Dict, List, Optional, Sequence, Tuple, Union + +import numpy as np + +from ..precision import Precision +from ..tensors import dtype_name, round_to_bf16, ttnn_dtype + +__all__ = [ + "CONV_PRECISION", + "FUSED_ACTIVATIONS", + "BINARY_FUSED_ACTIVATIONS", + "FeatureMap", + "Conv2d", + "KSplitConv", + "ConvTranspose2d", + "normalize_activation", + "unary_with_param", + "apply_activation", + "conv_padding", + "conv_out_hw", + "merge_sibling_convs", + "split_k_weights", + "concat_channels", + "split_channels", + "residual_add", + "fmap_from_numpy", + "fmap_to_numpy", + "is_dram", +] + +# Default precision of every builder here: HiFi4 + fp32 accumulation + packer L1 accumulation. Without +# ``packer_l1_acc`` a conv whose inner dim spans several blocks keeps its partial sums between blocks in the OUTPUT +# dtype (bf16) (conv2d_op_program_factory_common.cpp:175-178: partials CB = output dtype unless packer_l1_acc), which +# costs accuracy on wide convs; with it the partials stay fp32 (C17 device test, logs/ttaw/c17_c18_device_results.json). +CONV_PRECISION = "HiFi4+fp32+l1acc" + +# Activations a conv epilogue applies correctly on this tree (P5 / P14 on the device). HARDSWISH is NOT among them: +# a fused HARDSWISH compiles to nothing (unary_op_utils.cpp:964-965, probe P5) and returns the un-activated result. +FUSED_ACTIVATIONS = ("relu", "relu6", "silu", "gelu", "sigmoid") +# Activations fused into ``ttnn.add(..., activations=[...])`` verified on the device (P5: RELU6 on the clip6 add). +BINARY_FUSED_ACTIVATIONS = ("relu", "relu6") +_REFUSED = { + "hardswish": "a fused HARDSWISH is silently skipped by ttnn on this tree (probe P5): use activation=None and " + "ttnn.hardswish / ttnn.unary_chain([HARDSWISH]) as a separate op", +} + +ActivationSpec = Optional[str] + + +def normalize_activation(activation: Any) -> ActivationSpec: + """``None`` / ``"none"`` / ``"linear"`` / ``"identity"`` -> None; a fused activation name -> its lower-case name. + Raises ``ValueError`` for HARDSWISH (silently skipped when fused, P5) and for unknown names.""" + if activation is None: + return None + name = str(activation).strip().lower() + if name in ("", "none", "linear", "identity"): + return None + if name in _REFUSED: + raise ValueError(f"activation {activation!r}: {_REFUSED[name]}") + if name not in FUSED_ACTIVATIONS: + raise ValueError(f"activation {activation!r}: expected None or one of {FUSED_ACTIVATIONS}") + return name + + +def unary_with_param(activation: Any): + """The ``ttnn.UnaryWithParam`` of a fused activation (None for no activation). GELU is the exact (erf) form.""" + import ttnn + + name = normalize_activation(activation) + if name is None: + return None + if name == "gelu": + return ttnn.UnaryWithParam(ttnn.UnaryOpType.GELU, 0.0) + return ttnn.UnaryWithParam(getattr(ttnn.UnaryOpType, name.upper())) + + +def apply_activation(tensor: Any, activation: Any): + """The activation as a separate (unfused) ttnn op; identity for None.""" + import ttnn + + name = normalize_activation(activation) + if name is None: + return tensor + if name == "relu6": + return ttnn.relu6(tensor) + if name == "gelu": + return ttnn.gelu(tensor, fast_and_approximate_mode=False) + return getattr(ttnn, name)(tensor) + + +def _pair(value: Union[int, Sequence[int]], what: str) -> Tuple[int, int]: + if isinstance(value, (int, np.integer)): + return int(value), int(value) + v = tuple(int(x) for x in value) + if len(v) != 2: + raise ValueError(f"{what} must be an int or a pair, got {value!r}") + return v + + +def conv_padding(kernel: Tuple[int, int], padding: Any = "same", dilation: Tuple[int, int] = (1, 1) + ) -> Tuple[int, int, int, int]: + """Padding as (top, bottom, left, right). ``"same"``: ``dilation * (k // 2)`` on every side (odd kernels; the + "same" padding of stride-1 convs in every model here). An int, a pair (h, w) or a 4-tuple pass through.""" + if isinstance(padding, str): + if padding != "same": + raise ValueError(f"padding {padding!r}: expected 'same', an int, a pair or (top, bottom, left, right)") + if kernel[0] % 2 == 0 or kernel[1] % 2 == 0: + raise ValueError(f"'same' padding needs odd kernels, got {kernel}") + ph, pw = dilation[0] * (kernel[0] // 2), dilation[1] * (kernel[1] // 2) + return ph, ph, pw, pw + if isinstance(padding, (int, np.integer)): + p = int(padding) + return p, p, p, p + p = tuple(int(x) for x in padding) + if len(p) == 2: + return p[0], p[0], p[1], p[1] + if len(p) == 4: + return p # type: ignore[return-value] + raise ValueError(f"padding {padding!r}: expected 'same', an int, a pair or (top, bottom, left, right)") + + +def conv_out_hw(hw: Tuple[int, int], kernel: Tuple[int, int], stride: Tuple[int, int], + padding: Tuple[int, int, int, int], dilation: Tuple[int, int] = (1, 1)) -> Tuple[int, int]: + """Output size of a convolution (floor), as torch / ttnn compute it.""" + h = (hw[0] + padding[0] + padding[1] - dilation[0] * (kernel[0] - 1) - 1) // stride[0] + 1 + w = (hw[1] + padding[2] + padding[3] - dilation[1] * (kernel[1] - 1) - 1) // stride[1] + 1 + if h <= 0 or w <= 0: + raise ValueError(f"conv of a {hw} input with kernel {kernel}, stride {stride}, padding {padding} is empty") + return h, w + + +def is_dram(tensor_or_config: Any) -> bool: + """True when a device tensor (or a memory config) lives in DRAM.""" + import ttnn + + mc = tensor_or_config.memory_config() if hasattr(tensor_or_config, "memory_config") else tensor_or_config + bt = getattr(mc, "buffer_type", None) + if bt is not None and hasattr(ttnn, "BufferType"): + return bt == ttnn.BufferType.DRAM + return mc == ttnn.DRAM_MEMORY_CONFIG + + +def _memory_config(spec: Any): + """``"dram"`` / ``"l1"`` / None / a ttnn.MemoryConfig -> a memory config (or None).""" + import ttnn + + if spec is None: + return None + if isinstance(spec, str): + key = spec.lower() + if key == "dram": + return ttnn.DRAM_MEMORY_CONFIG + if key == "l1": + return ttnn.L1_MEMORY_CONFIG + raise ValueError(f"memory config {spec!r}: expected 'dram', 'l1', None or a ttnn.MemoryConfig") + return spec + + +def _layout(spec: Any): + import ttnn + + if isinstance(spec, str): + key = spec.lower().replace("_layout", "") + if key == "tile": + return ttnn.TILE_LAYOUT + if key in ("row_major", "rm"): + return ttnn.ROW_MAJOR_LAYOUT + raise ValueError(f"layout {spec!r}: expected 'tile' or 'row_major'") + return spec + + +def _layout_name(spec: Any) -> str: + """``"tile"`` / ``"row_major"`` of a layout spec (string or ttnn layout), without importing ttnn for strings.""" + if isinstance(spec, str): + key = spec.lower().replace("_layout", "") + return "row_major" if key in ("row_major", "rm") else key + return "tile" if "TILE" in str(spec).upper() else "row_major" + + +def _as_float32(a: Any, what: str) -> np.ndarray: + if hasattr(a, "detach"): + a = a.detach().cpu().float().numpy() + arr = np.asarray(a, dtype=np.float32) + if not np.all(np.isfinite(arr)): + raise ValueError(f"{what} holds non-finite values") + return np.ascontiguousarray(arr) + + +def _host_tensor(a: np.ndarray, dtype: str = "bfloat16"): + """A ROW_MAJOR ttnn *host* tensor (no device involved: only TILE host tensors open the device, probe G4-X). + The data are copied (ONNX initializers are read-only arrays, which torch warns about).""" + import torch + import ttnn + + return ttnn.from_torch(torch.from_numpy(np.array(a, dtype=np.float32, copy=True)), dtype=ttnn_dtype(dtype), + layout=ttnn.ROW_MAJOR_LAYOUT) + + +# --------------------------------------------------------------------------------------------- feature maps + +@dataclass(frozen=True) +class FeatureMap: + """A device feature map: ``tensor`` holds N*H*W rows of C channels (ttnn's flattened ``[1, 1, N*H*W, C]``, + TILE or ROW_MAJOR; a ``[N, H, W, C]`` ROW_MAJOR tensor is accepted too), plus the logical sizes ttnn's layout + does not keep.""" + + tensor: Any + batch: int + height: int + width: int + channels: int + + @property + def rows(self) -> int: + return self.batch * self.height * self.width + + @property + def hw(self) -> Tuple[int, int]: + return self.height, self.width + + def with_tensor(self, tensor: Any, channels: Optional[int] = None) -> "FeatureMap": + """The same spatial map with another tensor (and channel count).""" + return FeatureMap(tensor, self.batch, self.height, self.width, + self.channels if channels is None else int(channels)) + + def deallocate(self) -> None: + """Free the tensor's device memory (``force=False``: a view's base is left to its owner, probe P2).""" + import ttnn + + if self.tensor is not None and self.tensor.is_allocated(): + ttnn.deallocate(self.tensor, False) + + +def fmap_from_numpy(nchw: Any, device: Any, *, dtype: Any = "bfloat16", layout: Any = "tile", + memory_config: Any = "dram") -> FeatureMap: + """Upload an NCHW array as a device feature map ``[1, 1, N*H*W, C]`` (TILE by default). A TILE host tensor + touches the device: run under ``bin/devrun`` (probe G4-X).""" + import torch + import ttnn + + a = _as_float32(nchw, "feature map") if np.asarray(nchw).dtype != np.uint8 else np.asarray(nchw) + if a.ndim != 4: + raise ValueError(f"expected NCHW, got shape {a.shape}") + n, c, h, w = a.shape + flat = np.ascontiguousarray(a.transpose(0, 2, 3, 1).reshape(1, 1, n * h * w, c)) + t = ttnn.from_torch(torch.from_numpy(flat), dtype=ttnn_dtype(dtype), layout=_layout(layout), device=device, + memory_config=_memory_config(memory_config)) + return FeatureMap(t, n, h, w, c) + + +def fmap_to_numpy(fm: FeatureMap, dtype: Any = np.float32) -> np.ndarray: + """Read a device feature map back as an NCHW numpy array (blocking; never inside a capture).""" + import ttnn + + a = ttnn.to_torch(fm.tensor) + a = a.float().numpy() if a.is_floating_point() else a.numpy() + a = a.reshape(-1, a.shape[-1])[: fm.rows, : fm.channels] + return np.ascontiguousarray(a.reshape(fm.batch, fm.height, fm.width, fm.channels).transpose(0, 3, 1, 2), + dtype=dtype) + + +# ------------------------------------------------------------------------------------------------- Conv2d + +class Conv2d: + """``ttnn.conv2d`` with weights prepared by the first call and reused afterwards (see the module docstring). + + Args: + weight: float array ``[Cout, Cin / groups, kh, kw]`` (OIHW, numpy or torch; BN already folded). + bias: float array ``[Cout]`` or None. + stride, dilation: int or pair. ``padding``: ``"same"`` (default), int, pair or (top, bottom, left, right). + groups: 1 for dense convs; ``groups == Cin`` is a depthwise conv (correct but expanded to dense weights by + ttnn: P5, 311 MiB for SceneSeg). + activation: None or one of :data:`FUSED_ACTIVATIONS` (fused in the epilogue). + precision: a :class:`..precision.Precision` or its spec (default :data:`CONV_PRECISION` = HiFi4 + fp32 + accumulation + packer L1 accumulation; ``Precision.weights`` is the device weight dtype: bf16 or bfp8). + output_dtype: dtype of the output (``"bfloat16"``; ``"float32"`` for logits that feed thresholds). + output_layout: ``"tile"`` (default) or ``"row_major"``. + slicing: ``"auto"`` (ttnn's rule: DRAM input -> automatic DRAM slicing, L1 input -> L1 path), ``"l1"`` + (``Conv2dL1FullSliceConfig``: the op moves the whole input to L1) or ``("height" | "width", n)`` + (explicit DRAM slices; P14 found most explicit counts rejected or slower: diagnostics only). + output_memory_config: where an L1-path output (1x1 matmul convs, L1 inputs) is placed: ``"dram"`` + (default), ``"l1"``, None (the op's sharded output) or a ttnn.MemoryConfig. DRAM-sliced convs always + write DRAM interleaved; ``"l1"`` then adds a move. + conv_config: extra ``ttnn.Conv2dConfig`` fields (``act_block_h_override``, ``shard_layout``, + ``enable_act_double_buffer``, ``deallocate_activation``, ...): optimization knobs. + weight_terms: 1 (default) or 2. With 2 the weight and the bias are split into two bf16 terms + (``hi = bf16(w)``, ``lo = bf16(w - hi)``: about 16 significant bits) run as two convs into float32 + outputs, summed in fp32 with the activation fused into the add, then cast to ``output_dtype``: the conv + then carries ~fp32 weights and bias for twice the MACs and 3 more programs. The device's own float32 + weight path is no substitute (1.5x less error than bf16, vs 3.5x for the split; YOLOX port, + logs/yolox/m3_weight_precision_probe.log). Use it where the bf16 rounding of weights and biases, which + does not average out over a deep network (a per-channel bias error is the same at every pixel), + decides an output: YOLOX's 16-class mask on public frames (97.5 % -> 99.3 % agreement in emulation). + name: label for messages and :meth:`describe`. + """ + + def __init__(self, weight: Any, bias: Any = None, *, stride: Any = 1, padding: Any = "same", + dilation: Any = 1, groups: int = 1, activation: Any = None, precision: Any = CONV_PRECISION, + output_dtype: Any = "bfloat16", output_layout: Any = "tile", slicing: Any = "auto", + output_memory_config: Any = "dram", conv_config: Optional[Dict[str, Any]] = None, + weight_terms: int = 1, name: str = "conv"): + w = _as_float32(weight, f"{name}: weight") + if w.ndim != 4: + raise ValueError(f"{name}: weight must be [Cout, Cin/groups, kh, kw], got {w.shape}") + self.name = name + self.groups = int(groups) + self.out_channels, cin_g, kh, kw = (int(s) for s in w.shape) + self.in_channels = cin_g * self.groups + if self.groups < 1 or self.out_channels % self.groups: + raise ValueError(f"{name}: groups={groups} does not divide Cout={self.out_channels}") + self.kernel = (kh, kw) + self.stride = _pair(stride, "stride") + self.dilation = _pair(dilation, "dilation") + self.padding = conv_padding(self.kernel, padding, self.dilation) + self.activation = normalize_activation(activation) + self.precision = Precision.parse(precision) + self.output_dtype = dtype_name(output_dtype) + self.output_layout = output_layout + if not (slicing == "auto" or slicing == "l1" or (isinstance(slicing, (tuple, list)) and len(slicing) == 2 + and slicing[0] in ("height", "width"))): + raise ValueError(f"{name}: slicing {slicing!r}: expected 'auto', 'l1' or ('height' | 'width', n)") + self.slicing = slicing if isinstance(slicing, str) else (str(slicing[0]), int(slicing[1])) + self.output_memory_config = output_memory_config + self.conv_config_overrides = dict(conv_config or {}) + if "activation" in self.conv_config_overrides or "weights_dtype" in self.conv_config_overrides: + raise ValueError(f"{name}: pass activation / weights dtype as arguments, not in conv_config") + b = np.zeros(self.out_channels, np.float32) if bias is None else _as_float32(bias, f"{name}: bias") + if b.shape != (self.out_channels,): + raise ValueError(f"{name}: bias must have shape ({self.out_channels},), got {b.shape}") + if int(weight_terms) not in (1, 2): + raise ValueError(f"{name}: weight_terms must be 1 or 2, got {weight_terms!r}") + self.weight_terms = int(weight_terms) + if self.weight_terms == 2 and str(_layout_name(output_layout)) != "tile": + raise ValueError(f"{name}: weight_terms=2 writes TILE outputs (the fp32 sum is cast in TILE)") + self._weight = w + self._bias = b + self._host: Optional[List[Tuple[Any, Any]]] = None + self._prepared: Dict[Tuple, Tuple[Any, Any]] = {} + + # ---- geometry + @property + def is_matmul(self) -> bool: + """1x1, stride 1, no padding, no dilation, dense: ttnn runs it as a matmul on the L1 path.""" + return (self.kernel == (1, 1) and self.stride == (1, 1) and self.padding == (0, 0, 0, 0) + and self.dilation == (1, 1) and self.groups == 1) + + def out_hw(self, height: int, width: int) -> Tuple[int, int]: + return conv_out_hw((height, width), self.kernel, self.stride, self.padding, self.dilation) + + @property + def weight(self) -> np.ndarray: + """The float32 host weight (OIHW).""" + return self._weight + + @property + def bias(self) -> np.ndarray: + return self._bias + + # ---- ttnn arguments + def host_terms(self) -> List[Tuple[np.ndarray, np.ndarray]]: + """The float32 (weight, bias) of each term: one term, or ``hi = bf16(x)`` and ``lo = bf16(x - hi)`` (both + exactly representable in bf16, so the bf16 upload is exact).""" + if self.weight_terms == 1: + return [(self._weight, self._bias)] + hi_w, hi_b = round_to_bf16(self._weight), round_to_bf16(self._bias) + return [(hi_w, hi_b), (round_to_bf16(self._weight - hi_w), round_to_bf16(self._bias - hi_b))] + + def _host_tensors(self, term: int = 0) -> Tuple[Any, Any]: + if self._host is None: + # host tensors in the device weight dtype's host precision: bf16 for bf16 / bfp8 weights (ttnn rounds + # to bfp8 itself), float32 for float32 weights (a bf16 host tensor would round them first) + host_dtype = "float32" if self.precision.weights == "float32" else "bfloat16" + self._host = [(_host_tensor(w, host_dtype), _host_tensor(b.reshape(1, 1, 1, -1), host_dtype)) + for w, b in self.host_terms()] + return self._host[term] + + def _conv_config(self, fused: bool = True): + import ttnn + + kw = dict(self.conv_config_overrides) + kw["weights_dtype"] = ttnn_dtype(self.precision.weights) + kw["activation"] = unary_with_param(self.activation) if fused else None + kw.setdefault("output_layout", _layout(self.output_layout)) + return ttnn.Conv2dConfig(**kw) + + def _slice_config(self): + import ttnn + + if self.slicing == "auto": + return None + if self.slicing == "l1": + return ttnn.Conv2dL1FullSliceConfig + kind, num = self.slicing + slice_type = ttnn.Conv2dDRAMSliceHeight if kind == "height" else ttnn.Conv2dDRAMSliceWidth + return ttnn.Conv2dSliceConfig(slice_type=slice_type, num_slices=num) + + def _spec_key(self, fm: FeatureMap) -> Tuple: + t = fm.tensor + return (fm.batch, fm.height, fm.width, str(t.dtype), str(t.layout), str(t.memory_config()), + tuple(int(s) for s in t.shape)) + + def takes_l1_path(self, x: Any) -> bool: + """Whether a call on ``x`` (a device tensor) runs ttnn's L1 path (else the DRAM-sliced path).""" + if self.slicing == "l1" or self.is_matmul: + return True + return self.slicing == "auto" and not is_dram(x) + + def __call__(self, x: Union[FeatureMap, Any], *, batch: Optional[int] = None, height: Optional[int] = None, + width: Optional[int] = None) -> FeatureMap: + """Run the conv on a :class:`FeatureMap` (or a raw device tensor plus ``batch`` / ``height`` / ``width``) + and return the output :class:`FeatureMap` (``[1, 1, N*OH*OW, Cout]``).""" + import ttnn + + fm = x if isinstance(x, FeatureMap) else FeatureMap(x, int(batch), int(height), int(width), self.in_channels) + if fm.channels != self.in_channels: + raise ValueError(f"{self.name}: input has {fm.channels} channels, the conv takes {self.in_channels}") + key = self._spec_key(fm) + l1_path = self.takes_l1_path(fm.tensor) + out_mc = _memory_config(self.output_memory_config) + split = self.weight_terms == 2 + outs = [] + for term in range(self.weight_terms): + tkey = key + (term,) + weight, bias = self._prepared.get(tkey) or self._host_tensors(term) + out, (oh, ow), (weight, bias) = ttnn.conv2d( + input_tensor=fm.tensor, weight_tensor=weight, device=fm.tensor.device(), + in_channels=self.in_channels, out_channels=self.out_channels, batch_size=fm.batch, + input_height=fm.height, input_width=fm.width, kernel_size=self.kernel, stride=self.stride, + padding=self.padding, dilation=self.dilation, groups=self.groups, bias_tensor=bias, + conv_config=self._conv_config(fused=not split), compute_config=self.precision.compute_kernel_config(), + slice_config=self._slice_config(), memory_config=out_mc if l1_path else None, + dtype=ttnn.float32 if split else ttnn_dtype(self.output_dtype), return_output_dim=True, + return_weights_and_bias=True) + self._prepared[tkey] = (weight, bias) + outs.append(out) + if (int(oh), int(ow)) != self.out_hw(fm.height, fm.width): + raise RuntimeError(f"{self.name}: ttnn returned a {oh}x{ow} output, expected " + f"{self.out_hw(fm.height, fm.width)}") + out = outs[0] + if split: + fuse = self.activation in BINARY_FUSED_ACTIVATIONS + out = ttnn.add(outs[0], outs[1], activations=[unary_with_param(self.activation)] if fuse else []) + for t in outs: + ttnn.deallocate(t, False) + if not fuse: + out = apply_activation(out, self.activation) + if self.output_dtype != "float32": + out = ttnn.typecast(out, ttnn_dtype(self.output_dtype)) + if not l1_path and out_mc is not None and not is_dram(out_mc): + out = ttnn.to_memory_config(out, out_mc) + return FeatureMap(out, fm.batch, int(oh), int(ow), self.out_channels) + + @property + def prepared_specs(self) -> int: + """Number of input specs this layer holds prepared device weights for.""" + return len({k[:-1] for k in self._prepared}) + + def release(self) -> None: + """Free the prepared device weights (the host weights stay; the next call prepares them again).""" + import ttnn + + for w, b in self._prepared.values(): + for t in (w, b): + if t is not None and t.is_allocated(): + ttnn.deallocate(t) + self._prepared.clear() + + def describe(self) -> Dict[str, Any]: + return {"name": self.name, "cin": self.in_channels, "cout": self.out_channels, "kernel": list(self.kernel), + "stride": list(self.stride), "padding": list(self.padding), "dilation": list(self.dilation), + "groups": self.groups, "activation": self.activation, "precision": self.precision.label, + "output_dtype": self.output_dtype, "slicing": self.slicing if isinstance(self.slicing, str) + else list(self.slicing), "matmul": self.is_matmul, "weight_terms": self.weight_terms} + + +# ----------------------------------------------------------------------------------- K-split (concat-free) + +def split_k_weights(weight: Any, parts: Sequence[int]) -> List[np.ndarray]: + """Split a conv weight ``[Cout, sum(parts), kh, kw]`` along its input channels: ``conv(concat(x_i), W) == + sum_i conv(x_i, W_i)``.""" + w = _as_float32(weight, "weight") + if w.ndim != 4 or w.shape[1] != sum(int(p) for p in parts): + raise ValueError(f"weight {w.shape} does not have sum(parts) = {sum(parts)} input channels") + out, start = [], 0 + for p in parts: + out.append(np.ascontiguousarray(w[:, start:start + int(p)])) + start += int(p) + return out + + +class KSplitConv: + """``conv(concat([x_1, ..., x_n], channels))`` computed as ``sum_i conv_i(x_i)`` without building the concat + (S:centerpoint:448; P14: CenterPoint's shared conv 3.6 vs 14.8 ms and PCC 0.999992 vs 0.999961). + + The bias rides on the first partial conv, the partial sums are added in order, and the activation is fused into + the last add (RELU / RELU6, P5) or applied after it. ``accumulate_dtype="float32"`` keeps the partial outputs and + their sums in fp32 and casts to ``output_dtype`` once (more accurate, more traffic); the default ``"bfloat16"`` + rounds each partial (what P14 measured). Every keyword except ``activation`` / ``output_dtype`` / + ``accumulate_dtype`` goes to the partial :class:`Conv2d` layers.""" + + def __init__(self, weight: Any, bias: Any, parts: Sequence[int], *, activation: Any = None, + output_dtype: Any = "bfloat16", accumulate_dtype: Any = "bfloat16", name: str = "ksplit", + **conv_kwargs: Any): + if len(parts) < 2: + raise ValueError(f"{name}: a K-split needs at least two parts") + self.name = name + self.parts = tuple(int(p) for p in parts) + self.activation = normalize_activation(activation) + self.output_dtype = dtype_name(output_dtype) + self.accumulate_dtype = dtype_name(accumulate_dtype) + weights = split_k_weights(weight, self.parts) + cout = weights[0].shape[0] + b = np.zeros(cout, np.float32) if bias is None else _as_float32(bias, f"{name}: bias") + self.convs = [Conv2d(w, b if i == 0 else None, activation=None, output_dtype=self.accumulate_dtype, + name=f"{name}.{i}", **conv_kwargs) for i, w in enumerate(weights)] + self.out_channels = cout + + def __call__(self, inputs: Sequence[FeatureMap]) -> FeatureMap: + import ttnn + + if len(inputs) != len(self.convs): + raise ValueError(f"{self.name}: expected {len(self.convs)} inputs, got {len(inputs)}") + total: Optional[FeatureMap] = None + n = len(inputs) + for i, (conv, x) in enumerate(zip(self.convs, inputs)): + part = conv(x) + if total is None: + total = part + continue + last = i == n - 1 + fuse = last and self.activation in BINARY_FUSED_ACTIVATIONS + acts = [unary_with_param(self.activation)] if fuse else [] + summed = ttnn.add(total.tensor, part.tensor, activations=acts) + for t in (total, part): + t.deallocate() + total = total.with_tensor(summed) + if last and not fuse: + total = total.with_tensor(apply_activation(total.tensor, self.activation)) + assert total is not None + if self.accumulate_dtype != self.output_dtype: + total = total.with_tensor(ttnn.typecast(total.tensor, ttnn_dtype(self.output_dtype))) + return total + + def release(self) -> None: + for c in self.convs: + c.release() + + def describe(self) -> Dict[str, Any]: + return {"name": self.name, "parts": list(self.parts), "activation": self.activation, + "accumulate_dtype": self.accumulate_dtype, "convs": [c.describe() for c in self.convs]} + + +# ----------------------------------------------------------------------------------- sibling merge / glue + +def merge_sibling_convs(weights: Sequence[Any], biases: Sequence[Any]) -> Tuple[np.ndarray, np.ndarray, List[int]]: + """One conv for sibling convs reading the same input (S:centerpoint:450; S:yolox section 8.5 #1): output + channels concatenated in order -> ``(weight, bias, out_channels_per_sibling)``; split the output with + :func:`split_channels`. Exact (each output channel keeps its own weights).""" + ws = [_as_float32(w, "weight") for w in weights] + bs = [np.zeros(w.shape[0], np.float32) if b is None else _as_float32(b, "bias") for w, b in zip(ws, biases)] + shapes = {w.shape[1:] for w in ws} + if len(shapes) != 1: + raise ValueError(f"sibling convs must share Cin and kernel, got {[w.shape for w in ws]}") + return (np.ascontiguousarray(np.concatenate(ws, axis=0)), np.ascontiguousarray(np.concatenate(bs)), + [int(w.shape[0]) for w in ws]) + + +def concat_channels(maps: Sequence[FeatureMap], *, memory_config: Any = "dram") -> FeatureMap: + """Concatenate feature maps of one spatial size along channels (``ttnn.concat(dim=3)``); every input must have + the same layout and dtype. TILE inputs with channel counts that are multiples of 32 avoid any re-tiling.""" + import ttnn + + if not maps: + raise ValueError("concat_channels needs at least one map") + first = maps[0] + for m in maps[1:]: + if (m.batch, m.height, m.width) != (first.batch, first.height, first.width): + raise ValueError(f"concat of maps of different sizes: {[(m.batch, m.height, m.width) for m in maps]}") + if len(maps) == 1: + return first + t = ttnn.concat([m.tensor for m in maps], dim=3, memory_config=_memory_config(memory_config)) + return first.with_tensor(t, sum(m.channels for m in maps)) + + +def split_channels(fm: FeatureMap, sizes: Sequence[int]) -> List[FeatureMap]: + """Slice a map's channels into consecutive groups of ``sizes`` (``ttnn.slice`` on the last dim).""" + import ttnn + + if sum(int(s) for s in sizes) != fm.channels: + raise ValueError(f"split sizes {list(sizes)} do not add up to {fm.channels} channels") + shape = [int(s) for s in fm.tensor.shape] + out, start = [], 0 + for s in sizes: + begin = [0] * len(shape) + end = list(shape) + begin[-1], end[-1] = start, start + int(s) + out.append(fm.with_tensor(ttnn.slice(fm.tensor, begin, end), int(s))) + start += int(s) + return out + + +def residual_add(a: FeatureMap, b: FeatureMap, activation: Any = None) -> FeatureMap: + """``a + b`` (same spec), with RELU / RELU6 fused into the add (P5: the clip6 residual of YOLOX) or another + activation applied after it.""" + import ttnn + + if (a.rows, a.channels) != (b.rows, b.channels): + raise ValueError(f"residual add of {a.rows}x{a.channels} and {b.rows}x{b.channels}") + act = normalize_activation(activation) + if act in BINARY_FUSED_ACTIVATIONS: + return a.with_tensor(ttnn.add(a.tensor, b.tensor, activations=[unary_with_param(act)])) + return a.with_tensor(apply_activation(ttnn.add(a.tensor, b.tensor), act)) + + +# ----------------------------------------------------------------------------- ConvTranspose (kernel == stride) + +class ConvTranspose2d: + """ConvTranspose2d with kernel == stride and no padding (the BEV deblocks of CenterPoint / PointPainting / + TransFusion and SceneSeg's up-sampling layers), weights prepared by the first call. + + ``method="auto"`` follows probe P5: k = 2 -> ``ttnn.conv_transpose2d`` (CP deblock1 0.64 ms vs 2.08 ms for the + rewrite), k >= 3 -> ``"linear_d2s"`` = ``ttnn.linear(Cin -> k*k*Cout)`` with the bias and activation fused, then + depth-to-space (``to_layout(ROW_MAJOR)``, ``reshape [N*h, w, k, k, C]``, ``permute (0, 2, 1, 3, 4)``, ``reshape``; + CP deblock2 k = 4: 1.97 vs 10.27 ms). The linear output channel order is (i, j, c): + ``W2[cin, (i*k + j)*Cout + c] = W[cin, c, i, j]``. Output: a ROW_MAJOR ``[1, 1, N*h*k*w*k, Cout]`` map for + ``linear_d2s`` (``output_layout="tile"`` re-tiles it), the conv's TILE output for ``conv_transpose2d``. + + ``weight``: ``[Cin, Cout, k, k]`` (torch ConvTranspose2d layout); ``bias``: ``[Cout]`` or None.""" + + METHODS = ("auto", "conv_transpose2d", "linear_d2s") + + def __init__(self, weight: Any, bias: Any = None, *, stride: Optional[int] = None, activation: Any = None, + precision: Any = CONV_PRECISION, output_dtype: Any = "bfloat16", method: str = "auto", + output_layout: Any = "tile", memory_config: Any = "dram", name: str = "deconv"): + w = _as_float32(weight, f"{name}: weight") + if w.ndim != 4 or w.shape[2] != w.shape[3]: + raise ValueError(f"{name}: weight must be [Cin, Cout, k, k], got {w.shape}") + self.name = name + self.in_channels, self.out_channels, self.k = int(w.shape[0]), int(w.shape[1]), int(w.shape[2]) + if stride is not None and int(stride) != self.k: + raise ValueError(f"{name}: only kernel == stride is supported (k={self.k}, stride={stride})") + if method not in self.METHODS: + raise ValueError(f"{name}: method {method!r}: expected one of {self.METHODS}") + self.method = ("conv_transpose2d" if self.k == 2 else "linear_d2s") if method == "auto" else method + self.activation = normalize_activation(activation) + self.precision = Precision.parse(precision) + self.output_dtype = dtype_name(output_dtype) + self.output_layout = output_layout + self.memory_config = memory_config + b = np.zeros(self.out_channels, np.float32) if bias is None else _as_float32(bias, f"{name}: bias") + if b.shape != (self.out_channels,): + raise ValueError(f"{name}: bias must have shape ({self.out_channels},), got {b.shape}") + self._weight, self._bias = w, b + self._prepared: Dict[Tuple, Tuple[Any, Any]] = {} + self._linear: Optional[Tuple[Any, Any]] = None + + def _spec_key(self, fm: FeatureMap) -> Tuple: + t = fm.tensor + return (fm.batch, fm.height, fm.width, str(t.dtype), str(t.layout), str(t.memory_config())) + + def linear_weights(self) -> Tuple[np.ndarray, np.ndarray]: + """Host form of the ``linear_d2s`` weights: ``W2 [Cin, k*k*Cout]`` in (i, j, c) order, bias repeated.""" + k = self.k + w2 = self._weight.transpose(0, 2, 3, 1).reshape(self.in_channels, k * k * self.out_channels) + return np.ascontiguousarray(w2), np.ascontiguousarray(np.tile(self._bias, k * k)) + + def __call__(self, x: FeatureMap) -> FeatureMap: + import ttnn + + if x.channels != self.in_channels: + raise ValueError(f"{self.name}: input has {x.channels} channels, the layer takes {self.in_channels}") + k = self.k + if self.method == "conv_transpose2d": + key = self._spec_key(x) + # float32 host tensors for float32 weights, as Conv2d (a bf16 host tensor would round them first) + host_dtype = "float32" if self.precision.weights == "float32" else "bfloat16" + weight, bias = self._prepared.get(key) or (_host_tensor(self._weight, host_dtype), + _host_tensor(self._bias.reshape(1, 1, 1, -1), host_dtype)) + conv_config = ttnn.Conv2dConfig(weights_dtype=ttnn_dtype(self.precision.weights), + activation=unary_with_param(self.activation), + output_layout=_layout(self.output_layout)) + out, (weight, bias) = ttnn.conv_transpose2d( + input_tensor=x.tensor, weight_tensor=weight, device=x.tensor.device(), + in_channels=self.in_channels, out_channels=self.out_channels, batch_size=x.batch, + input_height=x.height, input_width=x.width, kernel_size=(k, k), stride=(k, k), padding=(0, 0), + output_padding=(0, 0), dilation=(1, 1), groups=1, bias_tensor=bias, conv_config=conv_config, + compute_config=self.precision.compute_kernel_config(), dtype=ttnn_dtype(self.output_dtype), + mirror_kernel=True, return_weights_and_bias=True) + self._prepared[key] = (weight, bias) + return FeatureMap(out, x.batch, x.height * k, x.width * k, self.out_channels) + # linear + depth-to-space + import torch + + device = x.tensor.device() + if self._linear is None: + w2, b2 = self.linear_weights() + wdt = ttnn_dtype(self.precision.weights) + self._linear = ( + ttnn.from_torch(torch.from_numpy(w2), dtype=wdt, layout=ttnn.TILE_LAYOUT, device=device), + ttnn.from_torch(torch.from_numpy(b2.reshape(1, -1)), dtype=wdt, layout=ttnn.TILE_LAYOUT, + device=device)) + weight, bias = self._linear + t = x.tensor if x.tensor.layout == ttnn.TILE_LAYOUT else ttnn.to_layout(x.tensor, ttnn.TILE_LAYOUT) + y = ttnn.linear(t, weight, bias=bias, activation=unary_with_param(self.activation), + compute_kernel_config=self.precision.compute_kernel_config(), + memory_config=_memory_config(self.memory_config), dtype=ttnn_dtype(self.output_dtype)) + y_rm = ttnn.to_layout(y, ttnn.ROW_MAJOR_LAYOUT) + ttnn.deallocate(y, False) + c = self.out_channels + y5 = ttnn.reshape(y_rm, (x.batch * x.height, x.width, k, k, c)) + y5 = ttnn.permute(y5, (0, 2, 1, 3, 4)) + out = ttnn.reshape(y5, (1, 1, x.batch * x.height * k * x.width * k, c)) + if _layout(self.output_layout) == ttnn.TILE_LAYOUT: + out = ttnn.to_layout(out, ttnn.TILE_LAYOUT) + return FeatureMap(out, x.batch, x.height * k, x.width * k, c) + + def release(self) -> None: + import ttnn + + tensors = [t for pair in self._prepared.values() for t in pair] + list(self._linear or ()) + for t in tensors: + if t is not None and t.is_allocated(): + ttnn.deallocate(t) + self._prepared.clear() + self._linear = None + + def describe(self) -> Dict[str, Any]: + return {"name": self.name, "cin": self.in_channels, "cout": self.out_channels, "k": self.k, + "method": self.method, "activation": self.activation, "precision": self.precision.label, + "output_dtype": self.output_dtype} diff --git a/code/tt_diffusion_planner/ttaw/ops/deform.py b/code/tt_diffusion_planner/ttaw/ops/deform.py new file mode 100644 index 0000000000000000000000000000000000000000..ecd110e51e6c380f500c86e381b93364e94ae3e4 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/ops/deform.py @@ -0,0 +1,182 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C23: ``ttnn.grid_sample`` helpers (PLAN.md section 1.2 row C23; probe P9, common/PROBES.md; first built by the +BEVDet port for AlignBEV and FPN_LSS). + +Rules from probe P9 (``common/probes/g4-attention-gridsample-msda.md``), enforced or encoded here: + +- **fp32 grids** for every data-dependent grid (bf16 grids cost PCC 0.9988 at 53 px and 0.992 at 128 px); +- **input C % 32 == 0** (C = 80 is rejected; zero pad channels stay exactly 0): :func:`padded_channels`, + :func:`pad_channels`; +- **per-batch grids**: grid N must equal input N (no broadcast over the batch); +- **ROW_MAJOR, interleaved (or height-sharded)** input and grid, ``padding_mode="zeros"`` only; +- **never a FLOAT32 input together with ``fp32_dest_acc_en``**: wrong, non-deterministic output (strict xfail in P9); + :func:`grid_sample` refuses it; +- **exact integer coordinates**: the kernel recovers pixel coordinates as ``x = g * W/2 + (W - 1)/2`` in fp32 + (``grid_sample_reader_common.hpp``, align_corners=False), so the align_corners=False normalisation + ``g = (2x + 1) / W - 1`` returns integer x exactly whenever ``(2x + 1) / W`` is exact in fp32 (``W`` a power of + two): an identity warp is then a bit-exact copy (P9: the literal align_corners=True normalisation leaves 2.8-7.6 % + of values off by up to 1/32). Zero padding acts in pixel space, so the normalisation choice does not change what + is sampled: callers state pixel coordinates (:func:`affine_grid`, :func:`resize_grid`) and always call the op with + ``align_corners=False``; +- the four bilinear weights are computed in fp32 and **truncated** (not rounded) to bf16 by the kernel: a 0.33-0.6 % + rrmse floor for fractional coordinates (plan risk R3; K5 with fp32 weights removes it). :func:`emulate_grid_sample` + reproduces that arithmetic on the host (the device test's oracle). + +Host functions are numpy only (grids are built on the host, per frame or once, and uploaded as fp32 ROW_MAJOR +tensors); :func:`grid_sample` is the device call with the checks above. DCNv2 as one grid_sample with K = 9 +(``batch_output_channels=True``, BEVFormer / METEOR) and the fused MSDA wrapper are added by their ports. +""" +from __future__ import annotations + +from typing import Any, Optional, Sequence, Tuple + +import numpy as np + +__all__ = ["OUTSIDE", "padded_channels", "pad_channels", "pixel_to_grid", "kernel_pixels", "affine_pixel_coords", + "affine_grid", "resize_pixel_coords", "resize_grid", "emulate_grid_sample", "grid_sample"] + +F32 = np.float32 +# A normalised coordinate whose four taps fall outside any input (pixel -1.5 W - 0.5): the sample is exactly 0. +OUTSIDE = F32(-4.0) + + +def padded_channels(c: int, multiple: int = 32) -> int: + """The channel count ``grid_sample`` accepts for ``c`` real channels (the next multiple of 32).""" + return -(-int(c) // int(multiple)) * int(multiple) + + +def pad_channels(x: np.ndarray, multiple: int = 32) -> np.ndarray: + """Zero-pad the last (channel) dimension of an NHWC array to a multiple of ``multiple`` (no copy when aligned).""" + x = np.asarray(x) + c = x.shape[-1] + cp = padded_channels(c, multiple) + if cp == c: + return x + out = np.zeros(x.shape[:-1] + (cp,), x.dtype) + out[..., :c] = x + return out + + +def pixel_to_grid(px: Any, size: int) -> np.ndarray: + """Pixel coordinate (0 = centre of the first pixel) -> normalised coordinate for ``align_corners=False``: + ``(2 px + 1) / size - 1`` in float32 (exact for integer ``px`` when ``size`` is a power of two).""" + p = np.asarray(px, F32) + return (F32(2.0) * p + F32(1.0)) / F32(size) - F32(1.0) + + +def kernel_pixels(g: Any, size: int) -> np.ndarray: + """The pixel coordinate the device kernel recovers from a normalised coordinate (align_corners=False): + ``g * size/2 + (size - 1)/2`` in float32.""" + g = np.asarray(g, F32) + return g * F32(size / 2.0) + F32((size - 1) / 2.0) + + +def affine_pixel_coords(transforms: np.ndarray, out_hw: Tuple[int, int]) -> Tuple[np.ndarray, np.ndarray]: + """Per batch n and output pixel (h, w): ``ix = a w + b h + c``, ``iy = d w + e h + f`` in float32 with one + rounding per operation (``transforms`` [N, 6] = (a, b, c, d, e, f), the BEVDet AlignBEV / BEVFormer-style pixel + affine) -> (ix, iy) [N, H, W].""" + t = np.asarray(transforms, F32).reshape(-1, 6) + h, w = out_hw + hh, ww = np.meshgrid(np.arange(h, dtype=F32), np.arange(w, dtype=F32), indexing="ij") + a, b, c, d, e, f = (t[:, k, None, None] for k in range(6)) + ix = (a * ww + b * hh) + c + iy = (d * ww + e * hh) + f + return ix.astype(F32), iy.astype(F32) + + +def affine_grid(transforms: np.ndarray, in_hw: Tuple[int, int], out_hw: Optional[Tuple[int, int]] = None, *, + valid: Optional[Sequence[bool]] = None) -> np.ndarray: + """fp32 grid [N, H, W, 2] (x, y) sampling an ``in_hw`` input at :func:`affine_pixel_coords` (align_corners=False + normalisation: identity transforms are bit-exact copies for power-of-two input sizes). Batches with + ``valid[n] == False`` sample :data:`OUTSIDE` (exact zeros, e.g. a history slot that holds no frame yet).""" + out_hw = tuple(in_hw) if out_hw is None else tuple(out_hw) + ix, iy = affine_pixel_coords(transforms, out_hw) + grid = np.stack([pixel_to_grid(ix, in_hw[1]), pixel_to_grid(iy, in_hw[0])], axis=-1).astype(F32) + if valid is not None: + v = np.asarray(valid, bool).reshape(-1) + if v.size != grid.shape[0]: + raise ValueError(f"valid has {v.size} entries for {grid.shape[0]} batches") + grid[~v] = OUTSIDE + return np.ascontiguousarray(grid) + + +def resize_pixel_coords(in_size: int, out_size: int, align_corners: bool = True) -> np.ndarray: + """Source pixel of every output pixel of a 1-D bilinear resize (float64), torch's ``F.interpolate`` rules: + align_corners=True ``i (in - 1) / (out - 1)``; False ``(i + 0.5) in / out - 0.5`` (not clamped: the zero padding + of grid_sample differs from torch's edge clamp there, so prefer C18 ``Resize2d`` for align_corners=False).""" + i = np.arange(out_size, dtype=np.float64) + if align_corners: + return i * (in_size - 1) / max(out_size - 1, 1) + return (i + 0.5) * in_size / out_size - 0.5 + + +def resize_grid(in_hw: Tuple[int, int], out_hw: Tuple[int, int], *, align_corners: bool = True, + batch: int = 1) -> np.ndarray: + """Static fp32 grid [batch, Ho, Wo, 2] of a bilinear resize (source pixels in float64, normalised with the + align_corners=False formula; call :func:`grid_sample`). For align_corners=True every source pixel lies inside + the input, so the result equals ``F.interpolate(..., align_corners=True)`` up to the bf16 weight truncation.""" + sy = resize_pixel_coords(in_hw[0], out_hw[0], align_corners) + sx = resize_pixel_coords(in_hw[1], out_hw[1], align_corners) + gx = ((2.0 * sx + 1.0) / in_hw[1] - 1.0).astype(F32) + gy = ((2.0 * sy + 1.0) / in_hw[0] - 1.0).astype(F32) + grid = np.empty((out_hw[0], out_hw[1], 2), F32) + grid[..., 0] = gx[None, :] + grid[..., 1] = gy[:, None] + return np.ascontiguousarray(np.broadcast_to(grid, (int(batch),) + grid.shape)) + + +def _truncate_bf16(a: np.ndarray) -> np.ndarray: + bits = np.ascontiguousarray(a, F32).view(np.uint32) & np.uint32(0xFFFF0000) + return bits.view(F32) + + +def emulate_grid_sample(x: np.ndarray, grid: np.ndarray) -> np.ndarray: + """Host emulation of the device's bilinear ``grid_sample`` (align_corners=False, zero padding): pixel recovery + as :func:`kernel_pixels`, ``floor``, fp32 fractions and weights truncated to bf16, taps outside the input skipped. + ``x`` [N, H, W, C], ``grid`` [N, Ho, Wo, 2] -> [N, Ho, Wo, C] float64 (the device sums the four bf16 products + in its own order, so compare with a tolerance; an exact identity / outside sample is exact on both).""" + x = np.asarray(x, np.float64) + g = np.asarray(grid, F32) + n, h, w, _ = x.shape + px = kernel_pixels(g[..., 0], w) + py = kernel_pixels(g[..., 1], h) + x0 = np.floor(px).astype(np.int64) + y0 = np.floor(py).astype(np.int64) + fx = (px - x0.astype(F32)).astype(F32) + fy = (py - y0.astype(F32)).astype(F32) + one = F32(1.0) + out = np.zeros(g.shape[:-1] + (x.shape[-1],), np.float64) + bidx = np.arange(n)[:, None, None] + for dy, wy in ((0, one - fy), (1, fy)): + for dx, wx in ((0, one - fx), (1, fx)): + yy, xx = y0 + dy, x0 + dx + ok = (yy >= 0) & (yy < h) & (xx >= 0) & (xx < w) + wgt = _truncate_bf16((wy * wx).astype(F32)).astype(np.float64) * ok + vals = x[bidx, np.clip(yy, 0, h - 1), np.clip(xx, 0, w - 1)] + out += vals * wgt[..., None] + return out + + +def grid_sample(x: Any, grid: Any, *, compute_kernel_config: Any = None, memory_config: Any = None, + batch_output_channels: bool = False): + """``ttnn.grid_sample(x, grid)``: bilinear, zero padding, ``align_corners=False`` (pair it with + :func:`pixel_to_grid` / :func:`affine_grid` / :func:`resize_grid`), with the P9 rules checked before the call. + The output goes to DRAM by default.""" + import ttnn + + if x.layout != ttnn.ROW_MAJOR_LAYOUT or grid.layout != ttnn.ROW_MAJOR_LAYOUT: + raise ValueError("grid_sample needs ROW_MAJOR input and grid") + if int(x.padded_shape[-1]) % 32: + raise ValueError(f"grid_sample input channels must be a multiple of 32, got {x.padded_shape[-1]} " + "(pad them: deform.pad_channels)") + if int(x.shape[0]) != int(grid.shape[0]): + raise ValueError(f"grid batch {grid.shape[0]} != input batch {x.shape[0]} (expand the grid per batch)") + if grid.dtype != ttnn.float32: + raise ValueError("use an fp32 grid for data-dependent sampling (probe P9: bf16 grids lose PCC)") + if x.dtype == ttnn.float32 and compute_kernel_config is not None and \ + getattr(compute_kernel_config, "fp32_dest_acc_en", False): + raise ValueError("a FLOAT32 input with fp32_dest_acc_en returns garbage (probe P9 defect)") + return ttnn.grid_sample(x, grid, mode="bilinear", padding_mode="zeros", align_corners=False, + use_precomputed_grid=False, batch_output_channels=batch_output_channels, + memory_config=ttnn.DRAM_MEMORY_CONFIG if memory_config is None else memory_config, + compute_kernel_config=compute_kernel_config) diff --git a/code/tt_diffusion_planner/ttaw/ops/gather.py b/code/tt_diffusion_planner/ttaw/ops/gather.py new file mode 100644 index 0000000000000000000000000000000000000000..6f32a6abca945e987405f2407350cc44764887b3 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/ops/gather.py @@ -0,0 +1,133 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C19: row gathers with ``ttnn.embedding`` and the gather form of scatter (PLAN.md section 1.2 row C19; probe P4). + +A scatter of ``P`` source rows into an ``N``-row destination (a BEV canvas, a to-dense map) is written as a +gather: the host builds ``index[n]`` = the source row of destination row ``n``, or ``sentinel`` (an explicit +all-zero row appended to the table) where nothing lands. Every shape is static, so it traces, and nothing has to +clear the destination first. Probe P4 (common/PROBES.md) measured it bit-exact at the real shapes, with: + +- the table ROW_MAJOR bf16 ``[1, 1, V, D]`` (``D % 32 == 0`` for the fused tilized output), UINT32 indices + ``[1, N]`` (ROW_MAJOR); +- **the default call** ``ttnn.embedding(index, table, layout=TILE, padding_idx=sentinel, + embeddings_type=PADDED)``: 0.19-0.54 ms per canvas (CenterPoint 480 x 480: 0.21 ms); GENERIC is 5.4-14.6x slower + on canvases (most indices hit the sentinel); +- **PADDED returns the table's own pad row, not zeros** (S:ptv3:304): the zero row must really be in the table + (:func:`with_zero_rows`). + +``ttnn.embedding`` is bf16-only; fp32 row gathers are ``ttnn.gather`` (bit-exact but ~1 us per index, probe P13), +not covered here. Indices must lie in ``[0, V)``: an out-of-range index reads outside the table (and can hang the +chip), so :func:`check_index` validates host index arrays before they are uploaded. + +numpy only at import; ttnn is imported inside the device functions. +""" +from __future__ import annotations + +from typing import Any, Optional, Tuple + +import numpy as np + +__all__ = [ + "check_index", + "gather_rows", + "gather_rows_numpy", + "scatter_rows", + "scatter_rows_numpy", + "with_zero_rows", +] + + +def check_index(index: Any, *, num_rows: int, sentinel: Optional[int] = None) -> np.ndarray: + """Validate a host gather index and return it as a contiguous ``uint32`` array of shape ``[1, N]``. + + ``num_rows`` is the number of real source rows; valid entries are ``0 <= i < num_rows`` or ``i == sentinel`` + (the zero row, conventionally ``num_rows``). Raises ``ValueError`` for anything else (negative, too large, + non-integral, non-finite), because an out-of-range index makes the device read outside the table.""" + a = np.asarray(index) + if a.dtype.kind == "f": + if not np.all(np.isfinite(a)) or np.any(a != np.floor(a)): + raise ValueError("gather index holds non-integral or non-finite values") + elif a.dtype.kind not in "iu": + raise ValueError(f"gather index must be integral, got dtype {a.dtype}") + flat = a.reshape(-1).astype(np.int64) + valid = (flat >= 0) & (flat < int(num_rows)) + if sentinel is not None: + valid |= flat == int(sentinel) + bad = ~valid + if np.any(bad): + i = int(np.flatnonzero(bad)[0]) + raise ValueError(f"gather index [{i}] = {int(flat[i])} is outside [0, {num_rows})" + + ("" if sentinel is None else f" and is not the sentinel {sentinel}")) + return np.ascontiguousarray(flat.astype(np.uint32).reshape(1, -1)) + + +def gather_rows_numpy(index: Any, table: Any) -> np.ndarray: + """Host oracle: ``out[n] = table[index[n]]`` (``table`` ``[V, D]``).""" + t = np.asarray(table) + return t.reshape(-1, t.shape[-1])[np.asarray(index).reshape(-1).astype(np.int64)] + + +def scatter_rows_numpy(rows: Any, index: Any, *, sentinel: int) -> np.ndarray: + """Host oracle of :func:`scatter_rows`: ``rows`` ``[P, D]`` with a zero row appended at ``sentinel`` + (``sentinel >= P``; rows between ``P`` and ``sentinel`` are zero too), gathered by ``index``.""" + r = np.asarray(rows, dtype=np.float32) + r = r.reshape(-1, r.shape[-1]) + table = np.zeros((max(int(sentinel) + 1, r.shape[0]), r.shape[1]), np.float32) + table[: r.shape[0]] = r + return gather_rows_numpy(index, table) + + +def _rm(t: Any): + import ttnn + + return ttnn.to_layout(t, ttnn.ROW_MAJOR_LAYOUT) if t.layout != ttnn.ROW_MAJOR_LAYOUT else t + + +def with_zero_rows(rows: Any, count: int = 1): + """A ROW_MAJOR gather table ``[1, 1, P + count, D]``: the device rows ``rows`` (``[1, 1, P, D]``, TILE or + ROW_MAJOR, bf16) followed by ``count`` zero rows (the sentinel rows of :func:`scatter_rows`). One ``ttnn.pad`` + on the device (plus an untilize for a TILE input); traceable.""" + import ttnn + + if count < 1: + raise ValueError("count must be >= 1") + shape = tuple(int(s) for s in rows.shape) + if len(shape) != 4 or shape[0] != 1 or shape[1] != 1: + raise ValueError(f"rows must be [1, 1, P, D], got {shape}") + rm = _rm(rows) + return ttnn.pad(rm, [(0, 0), (0, 0), (0, int(count)), (0, 0)], 0.0) + + +def gather_rows(index: Any, table: Any, *, sentinel: Optional[int] = None, layout: str = "tile", + memory_config: Any = None): + """``out[0, 0, n] = table[0, 0, index[0, n]]`` on the device -> ``[1, 1, N, D]``. + + ``index``: UINT32 ROW_MAJOR ``[1, N]`` device tensor; ``table``: bf16 ROW_MAJOR ``[1, 1, V, D]`` (a TILE table + is untilized by ttnn first: one extra program). ``sentinel``: the index of an all-zero table row; given, the + call is P4's fast PADDED form (the table must really hold zeros there). ``layout="tile"`` (default) writes + the tilized output directly when ``N % 32 == 0`` and ``D % 32 == 0``.""" + import ttnn + + lay = ttnn.TILE_LAYOUT if layout == "tile" else ttnn.ROW_MAJOR_LAYOUT + kw = {"layout": lay} + if sentinel is not None: + kw.update(padding_idx=int(sentinel), embeddings_type=ttnn.EmbeddingsType.PADDED) + if memory_config is not None: + kw["memory_config"] = memory_config + out = ttnn.embedding(index, table, **kw) # [1, N, D] + n, d = int(out.shape[-2]), int(out.shape[-1]) + return ttnn.reshape(out, (1, 1, n, d)) + + +def scatter_rows(rows: Any, index: Any, *, sentinel: Optional[int] = None, layout: str = "tile", + memory_config: Any = None) -> Tuple[Any, Any]: + """Gather-form scatter: ``canvas[n] = rows[index[n]]``, or zeros where ``index[n] == sentinel``. + + ``rows``: ``[1, 1, P, D]`` device rows (TILE or ROW_MAJOR bf16); ``sentinel`` defaults to ``P`` (the zero + row appended by :func:`with_zero_rows`). Returns ``(canvas, table)``: the ``[1, 1, N, D]`` canvas (TILE by + default, e.g. the NHWC input of a conv) and the table, which the caller may deallocate.""" + p = int(rows.shape[-2]) + sentinel = p if sentinel is None else int(sentinel) + if sentinel < p: + raise ValueError(f"sentinel {sentinel} must be >= the number of source rows {p}") + table = with_zero_rows(rows, sentinel - p + 1) + return gather_rows(index, table, sentinel=sentinel, layout=layout, memory_config=memory_config), table diff --git a/code/tt_diffusion_planner/ttaw/ops/heatmap.py b/code/tt_diffusion_planner/ttaw/ops/heatmap.py new file mode 100644 index 0000000000000000000000000000000000000000..c331ea000033e57e8be553c508de5b46478193f2 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/ops/heatmap.py @@ -0,0 +1,245 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C22: heatmap peaks of the query-based LiDAR detectors -- sigmoid and the 3x3 local maximum (PLAN.md section 1.2 +row C22; probe P6, common/PROBES.md). + +TransFusion, BEVFusion and PTv3 keep a heatmap cell only where it is the maximum of its 3x3 neighbourhood, with +model-specific border and class rules: + +- ``"transfusion"`` (TF, S:transfusion:169): ``MaxPool 3x3 s1 p0`` padded back with a ZERO border, every class: the + border cells are always suppressed; +- ``"bevfusion"`` (BF, S:bevfusion:258): classes ``pooled_classes`` (default 0-3) as TF; the other classes are not + pooled (``local_max = heat``): every cell passes; +- ``"ptv3"`` (PT, S:ptv3:217): ``local_max`` starts as the heat itself and only the interior of the pooled classes is + replaced by the pool: border cells and the other classes pass. + +All three are one exact masked chain (P6, bit-exact on the device, TF 95 / BF 85 / PT 145 us): + + keep = m_pass + (heat == max_pool2d(heat, 3x3, stride 1, padding 1)) * m_eq out = heat * keep + +with two constant masks ``m_eq`` / ``m_pass`` (disjoint, values 0 / 1, laid out as the NHWC rows ``[H*W, C]`` of the +heatmap). The pool's padding never matters: a padded window only covers border cells, where ``m_eq`` is 0. Not the +literal ``p0`` pool + pad form (222 us for TF and two more programs). ``ttnn.max_pool2d`` is bf16 only (FLOAT32 is +rejected, ``pool_op.cpp:30``), so the heat is bf16: :func:`sigmoid_heat` takes fp32 logits through an fp32 sigmoid and +rounds once. + +Device layout: the heatmap is a TILE ``[1, 1, H*W, C]`` tensor (what a conv head writes); the result has the same +layout (the input of C21's class-major top-k row and of the query-heat gather of C26). + +numpy only at import; ttnn is imported inside the device functions. +""" +from __future__ import annotations + +from typing import Any, Dict, Optional, Sequence, Tuple + +import numpy as np + +__all__ = ["RULES", "DEFAULT_POOLED_CLASSES", "local_max_masks", "local_max_numpy", "local_max_literal_numpy", + "sigmoid_heat", "sigmoid_heat_numpy", "LocalMax"] + +RULES = ("transfusion", "bevfusion", "ptv3") +DEFAULT_POOLED_CLASSES = 4 # BF / PT: classes 0-3 pooled (the last class(es) pass) + + +def _pooled(classes: int, rule: str, pooled_classes: Optional[Sequence[int]]) -> np.ndarray: + if rule not in RULES: + raise ValueError(f"rule {rule!r}: expected one of {RULES}") + if pooled_classes is None: + pooled_classes = range(classes) if rule == "transfusion" else range(min(DEFAULT_POOLED_CLASSES, classes)) + p = np.zeros(classes, np.float32) + for c in pooled_classes: + if not 0 <= int(c) < classes: + raise ValueError(f"pooled class {c} outside [0, {classes})") + p[int(c)] = 1.0 + if rule == "transfusion" and not p.all(): + raise ValueError("the transfusion rule pools every class") + return p + + +def local_max_masks(height: int, width: int, classes: int, *, rule: str = "transfusion", + pooled_classes: Optional[Sequence[int]] = None, kernel: int = 3) -> Tuple[np.ndarray, np.ndarray]: + """The constant masks ``(m_pass, m_eq)`` of the chain for ``rule`` (module docstring), float32 NHWC rows + ``[height * width, classes]``. ``kernel`` (odd): the pool size; the border is ``kernel // 2`` cells wide.""" + if kernel < 1 or kernel % 2 == 0: + raise ValueError(f"kernel must be odd and positive, got {kernel}") + r = kernel // 2 + if height <= 2 * r or width <= 2 * r: + raise ValueError(f"a {height}x{width} map has no interior for a {kernel}x{kernel} pool") + pooled = _pooled(classes, rule, pooled_classes) + interior = np.zeros((height, width, 1), np.float32) + interior[r:height - r, r:width - r] = 1.0 + m_eq = interior * pooled[None, None, :] + if rule == "transfusion": + m_pass = np.zeros_like(m_eq) + elif rule == "bevfusion": + m_pass = np.broadcast_to(1.0 - pooled[None, None, :], m_eq.shape).astype(np.float32) + else: + m_pass = 1.0 - m_eq + return (np.ascontiguousarray(m_pass.reshape(height * width, classes)), + np.ascontiguousarray(m_eq.reshape(height * width, classes))) + + +def _pool_numpy(heat_nhwc: np.ndarray, height: int, width: int, kernel: int) -> np.ndarray: + """3x3 (k x k) stride-1 max pool with ``-inf`` padding of NHWC rows ``[H*W, C]`` (exact on any float data).""" + c = heat_nhwc.shape[-1] + x = np.asarray(heat_nhwc, np.float64).reshape(height, width, c) + r = kernel // 2 + p = np.full((height + 2 * r, width + 2 * r, c), -np.inf) + p[r:r + height, r:r + width] = x + out = np.full_like(x, -np.inf) + for dy in range(kernel): + for dx in range(kernel): + np.maximum(out, p[dy:dy + height, dx:dx + width], out=out) + return out.reshape(height * width, c) + + +def local_max_numpy(heat_nhwc: Any, height: int, width: int, m_pass: Any, m_eq: Any, kernel: int = 3) -> np.ndarray: + """Host oracle of the masked chain on NHWC rows ``[H*W, C]`` (float32 out; exact on bf16-valued input).""" + h = np.asarray(heat_nhwc, np.float32).reshape(height * width, -1) + pooled = _pool_numpy(h, height, width, kernel) + keep = np.asarray(m_pass, np.float32) + (h == pooled).astype(np.float32) * np.asarray(m_eq, np.float32) + return (h * keep).astype(np.float32) + + +def local_max_literal_numpy(heat_nchw: Any, rule: str = "transfusion", *, + pooled_classes: Optional[Sequence[int]] = None, kernel: int = 3) -> np.ndarray: + """The models' literal forms (``heat * (heat == local_max)``, NCHW ``(C, H, W)`` or ``(1, C, H, W)``), the oracle + the masked chain is proven against: TF ``max_pool p0`` + zero border (all classes); BF pooled classes as TF and + the others ``local_max = heat``; PT ``local_max = heat`` with the interior of the pooled classes pooled.""" + x = np.asarray(heat_nchw, np.float32) + squeeze = x.ndim == 4 + if squeeze: + x = x[0] + c, h, w = x.shape + pooled = _pooled(c, rule, pooled_classes).astype(bool) + r = kernel // 2 + full = _pool_numpy(x.reshape(c, h * w).T, h, w, kernel).T.reshape(c, h, w) + interior = np.zeros((h, w), bool) + interior[r:h - r, r:w - r] = True + if rule == "transfusion": + local = np.where(interior[None], full, 0.0) + elif rule == "bevfusion": + local = np.where(pooled[:, None, None], np.where(interior[None], full, 0.0), x) + else: + local = np.where(pooled[:, None, None] & interior[None], full, x) + out = (x * (x == local)).astype(np.float32) + return out[None] if squeeze else out + + +def sigmoid_heat_numpy(logits: Any, dtype: str = "bfloat16") -> np.ndarray: + """Host model of :func:`sigmoid_heat`: float64 sigmoid of the float32 logits, rounded once to ``dtype`` (bf16 RNE).""" + s = (1.0 / (1.0 + np.exp(-np.asarray(logits, np.float64)))).astype(np.float32) + if dtype in ("bfloat16", "bf16"): + from ..tensors import round_to_bf16 + + return round_to_bf16(s) + return s + + +def sigmoid_heat(logits: Any, *, dtype: str = "bfloat16", memory_config: Any = None): + """Device ``sigmoid(logits)`` computed in the logits' dtype (fp32 logits: an fp32 sigmoid) and cast once to + ``dtype`` (bf16: what ``max_pool2d`` and C21's ``topk_large_indices`` take). Returns ``(heat, sig)``: the cast + heat and the sigmoid before the cast (the same tensor when no cast is needed).""" + import ttnn + + from ..tensors import ttnn_dtype + + kw = {} if memory_config is None else {"memory_config": memory_config} + sig = ttnn.sigmoid(logits, **kw) + want = ttnn_dtype(dtype) + heat = sig if sig.dtype == want else ttnn.typecast(sig, want) + return heat, sig + + +def _mc(spec: Any): + import ttnn + + if spec is None or spec == "dram": + return ttnn.DRAM_MEMORY_CONFIG + if spec == "l1": + return ttnn.L1_MEMORY_CONFIG + return spec + + +class LocalMax: + """The device chain for one heatmap geometry: ``heat`` (bf16 TILE ``[1, 1, H*W, C]``) -> ``heat * keep``. + + The masks are uploaded (bf16 TILE, exact 0 / 1) by the first call, outside any capture; with the TF rule + (``m_pass`` = 0) the chain is ``max_pool2d``, ``eq``, ``multiply`` (keep = eq * m_eq), ``multiply``; otherwise + ``addcmul(m_pass, eq, m_eq)`` replaces the first multiply. ``memory_config``: where the pool and the masks live + (``"dram"`` default; ``"l1"`` is P6's 95 us form for maps that fit).""" + + def __init__(self, height: int, width: int, classes: int, *, rule: str = "transfusion", + pooled_classes: Optional[Sequence[int]] = None, kernel: int = 3, memory_config: Any = "dram", + name: str = "local_max"): + self.height, self.width, self.classes = int(height), int(width), int(classes) + self.rule = rule + self.kernel = int(kernel) + self.name = name + self.m_pass, self.m_eq = local_max_masks(self.height, self.width, self.classes, rule=rule, + pooled_classes=pooled_classes, kernel=self.kernel) + self.passes = bool(self.m_pass.any()) + self.memory_config = memory_config + self._dev: Optional[Dict[str, Any]] = None + + @property + def cells(self) -> int: + return self.height * self.width + + def _upload(self, device: Any) -> Dict[str, Any]: + import torch + import ttnn + + def up(a: np.ndarray): + t = torch.from_numpy(np.ascontiguousarray(a.reshape(1, 1, self.cells, self.classes))) + return ttnn.from_torch(t, dtype=ttnn.bfloat16, layout=ttnn.TILE_LAYOUT, device=device, + memory_config=_mc(self.memory_config)) + + self._dev = {"m_eq": up(self.m_eq), "m_pass": up(self.m_pass) if self.passes else None} + return self._dev + + def __call__(self, heat: Any): + import ttnn + + shape = tuple(int(s) for s in heat.shape) + if shape != (1, 1, self.cells, self.classes): + raise ValueError(f"{self.name}: heat must be [1, 1, {self.cells}, {self.classes}], got {shape}") + if heat.dtype != ttnn.bfloat16: + raise ValueError(f"{self.name}: max_pool2d takes bf16 only (got {heat.dtype}): use sigmoid_heat") + dev = self._dev or self._upload(heat.device()) + r = self.kernel // 2 + pooled = ttnn.max_pool2d(input_tensor=heat, batch_size=1, input_h=self.height, input_w=self.width, + channels=self.classes, kernel_size=[self.kernel, self.kernel], stride=[1, 1], + padding=[r, r], dilation=[1, 1], ceil_mode=False, + memory_config=_mc(self.memory_config), output_layout=ttnn.TILE_LAYOUT) + if tuple(int(s) for s in pooled.shape) != shape: + full = pooled + pooled = ttnn.slice(full, [0, 0, 0, 0], list(shape)) + ttnn.deallocate(full, False) + eq = ttnn.eq(heat, pooled) + ttnn.deallocate(pooled, False) + if self.passes: + keep = ttnn.addcmul(dev["m_pass"], eq, dev["m_eq"]) + else: + keep = ttnn.multiply(eq, dev["m_eq"]) + ttnn.deallocate(eq, False) + out = ttnn.multiply(heat, keep) + ttnn.deallocate(keep, False) + return out + + def numpy(self, heat_nhwc: Any) -> np.ndarray: + """Host oracle of this chain on NHWC rows (bf16-valued input gives the device result bit for bit).""" + return local_max_numpy(heat_nhwc, self.height, self.width, self.m_pass, self.m_eq, self.kernel) + + def release(self) -> None: + import ttnn + + if self._dev is not None: + for t in self._dev.values(): + if t is not None and t.is_allocated(): + ttnn.deallocate(t) + self._dev = None + + def describe(self) -> Dict[str, Any]: + return {"name": self.name, "grid": [self.height, self.width], "classes": self.classes, "rule": self.rule, + "kernel": self.kernel, "pooled_classes": [int(c) for c in np.flatnonzero(self.m_eq.max(axis=0))], + "passes": self.passes, "memory_config": str(self.memory_config)} diff --git a/code/tt_diffusion_planner/ttaw/ops/kernels/segment_reduce_dm.cpp b/code/tt_diffusion_planner/ttaw/ops/kernels/segment_reduce_dm.cpp new file mode 100644 index 0000000000000000000000000000000000000000..92ac7c19248c60c37e73cd3b3a193196485e45f4 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/ops/kernels/segment_reduce_dm.cpp @@ -0,0 +1,292 @@ +// SPDX-License-Identifier: Apache-2.0 +// K1 segment_reduce (ttaw.ops.segment, PLAN.md 1.3): per-segment MAX over the rows of a tensor sorted by segment. +// +// out[s, :] = max(x[off[s] : off[s + 1], :]) for s in [0, NUM_SEG); empty segments -> FILL_BITS +// +// x: [1, 1, NUM_ROWS, C] interleaved DRAM, TILE (IN_TILE = 1) or ROW_MAJOR, bf16 or fp32; off: uint32 ROW_MAJOR +// [1, >= NUM_SEG + 1] interleaved DRAM (one page; CSR offsets, non-decreasing); out: [1, 1, NUM_SEG, C] ROW_MAJOR +// interleaved DRAM, bf16 (round to nearest even) or fp32. Exact: the maximum is taken on integer keys of the +// float bit patterns (key = bits ^ ((bits >> 31) & 0x7fffffff), signed order == float order, -0 < +0), so no float +// arithmetic happens on the RISC-V (no F extension) and bf16 / fp32 inputs come out bit-exact (fp32 -> bf16 is a +// monotonic rounding, so max-then-round == round-then-max). NaN is not supported (not produced by the models). +// +// Work split (deterministic, data-dependent, no per-core runtime args): NUM_WORKERS workers = SLOTS data-movement +// RISC-Vs per core on a GW-wide core rectangle starting at (X0, Y0); worker w = core_index * SLOTS + SLOT. The cost +// of segment s is off[s] + s (its rows plus one output row; strictly increasing in s), T = off[S] + S. Worker w owns +// the segments whose cost lies in [w T / W, (w + 1) T / W): it finds them by two binary searches over off (64-byte +// probes from DRAM), then streams its contiguous row range in units of 32 rows (one tile-row in TILE mode) with +// double buffering, reduces each segment into an L1 accumulator and writes one output row per segment (a ring of +// OUT_RING staging rows). Every segment is written by exactly one worker; nothing is shared between workers. +// +// Runtime args: COMMON only [x_addr, off_addr, out_addr] (P3 rule: per-core work derives from the core coordinates; +// the arg layout is fixed by the compile-time args). Hang-protocol notes (PLAN.md 4.4): no CB is pushed or popped +// (CBs are plain L1 scratch: data 2 units, offsets window, probe block, accumulator, output ring; sizes checked by +// the host builder), no multicast, no semaphores; every DRAM read has matching 64-byte alignment of source and +// destination (Blackhole NOC_DRAM_READ_ALIGNMENT_BYTES), every read is waited on by noc_async_read_barrier (which +// also invalidates the Blackhole L1 data cache) before the data are used, staging rows are reused only after +// noc_async_writes_flushed, and the kernel ends with noc_async_write_barrier. Offsets are clamped to [0, NUM_ROWS] +// and to non-decreasing order, so a corrupt table can never make the kernel read outside x. +#include + +#include "api/dataflow/dataflow_api.h" + +constexpr uint32_t IN_TILE = get_compile_time_arg_val(0); // 1: TILE input, 0: ROW_MAJOR input +constexpr uint32_t IN_BYTES = get_compile_time_arg_val(1); // input element bytes: 2 (bf16) or 4 (fp32) +constexpr uint32_t OUT_BYTES = get_compile_time_arg_val(2); // output element bytes: 2 (bf16, RNE) or 4 (fp32) +constexpr uint32_t C = get_compile_time_arg_val(3); // channels (multiple of 32) +constexpr uint32_t NUM_ROWS = get_compile_time_arg_val(4); // input rows (multiple of 32): the read bound +constexpr uint32_t NUM_SEG = get_compile_time_arg_val(5); // output rows (segments) +constexpr uint32_t NUM_WORKERS = get_compile_time_arg_val(6); // workers of the program +constexpr uint32_t SLOTS = get_compile_time_arg_val(7); // workers per core (1 or 2) +constexpr uint32_t SLOT = get_compile_time_arg_val(8); // this kernel's slot on its core +constexpr uint32_t X0 = get_compile_time_arg_val(9); // core rectangle: first logical x +constexpr uint32_t Y0 = get_compile_time_arg_val(10); // first logical y +constexpr uint32_t GW = get_compile_time_arg_val(11); // width in cores +constexpr uint32_t CB_DATA = get_compile_time_arg_val(12); // 2 units of input rows (+ 64 B alignment slack) +constexpr uint32_t CB_OFF = get_compile_time_arg_val(13); // offsets window: OFF_CHUNK uint32 (+ 64 B) +constexpr uint32_t CB_PROBE = get_compile_time_arg_val(14); // one 64-byte probe block (+ 64 B) +constexpr uint32_t CB_ACC = get_compile_time_arg_val(15); // C int32 keys +constexpr uint32_t CB_OUT = get_compile_time_arg_val(16); // OUT_RING output rows (+ 64 B) +constexpr uint32_t OUT_RING = get_compile_time_arg_val(17); // staging rows in flight +constexpr uint32_t OFF_CHUNK = get_compile_time_arg_val(18); // offsets per window (multiple of 16) +constexpr uint32_t FILL_BITS = get_compile_time_arg_val(19); // fp32 bits written for an empty segment +constexpr uint32_t OFF_PAGE_BYTES = get_compile_time_arg_val(20); // page (= tensor) size of off, bytes +constexpr auto x_args = TensorAccessorArgs<21>(); +constexpr auto off_args = TensorAccessorArgs(); +constexpr auto out_args = TensorAccessorArgs(); + +constexpr uint32_t UNIT_ROWS = 32; // rows per streamed unit (one tile-row) +constexpr uint32_t CT = C / 32; // 32-column tiles per row +constexpr uint32_t TILE_ELEMS = 1024; +constexpr uint32_t IN_PAGE_BYTES = IN_TILE ? TILE_ELEMS * IN_BYTES : C * IN_BYTES; +constexpr uint32_t PAGES_PER_UNIT = IN_TILE ? CT : UNIT_ROWS; +constexpr uint32_t UNIT_BYTES = PAGES_PER_UNIT * IN_PAGE_BYTES; +constexpr uint32_t OUT_ROW_BYTES = C * OUT_BYTES; +constexpr uint32_t OFF_ALIGNED_BYTES = (OFF_PAGE_BYTES + 63u) & ~63u; +constexpr uint32_t NUM_UNITS = NUM_ROWS / UNIT_ROWS; + +static_assert(C % 32 == 0, "C must be a multiple of 32"); +static_assert(NUM_ROWS % UNIT_ROWS == 0, "NUM_ROWS must be a multiple of 32"); +static_assert(IN_BYTES == 2 || IN_BYTES == 4, "bf16 or fp32 input"); +static_assert(OUT_BYTES == 2 || OUT_BYTES == 4, "bf16 or fp32 output"); +static_assert(OFF_CHUNK % 16 == 0 && OFF_CHUNK >= 16, "OFF_CHUNK must be a multiple of 16"); +static_assert(OUT_RING >= 1, "OUT_RING >= 1"); +static_assert(IN_PAGE_BYTES % 64 == 0, "input pages must be 64-byte multiples (DRAM read alignment)"); +static_assert(OUT_ROW_BYTES % 16 == 0, "output rows must be 16-byte multiples (DRAM write alignment)"); + +#define COMPILER_BARRIER() asm volatile("" ::: "memory") + +static inline uint32_t align64(uint32_t a) { return (a + 63u) & ~63u; } + +// float bits <-> signed key (an involution); signed comparison of keys == float comparison (no NaN) +static inline int32_t to_key(uint32_t bits) { + const int32_t b = static_cast(bits); + return b ^ ((b >> 31) & 0x7fffffff); +} +static inline uint32_t from_key(int32_t k) { return static_cast(k ^ ((k >> 31) & 0x7fffffff)); } + +static inline uint16_t to_bf16_rne(uint32_t bits) { + return static_cast((bits + 0x7fffu + ((bits >> 16) & 1u)) >> 16); +} + +void kernel_main() { + const uint32_t x_addr = get_common_arg_val(0); + const uint32_t off_addr = get_common_arg_val(1); + const uint32_t out_addr = get_common_arg_val(2); + + const uint32_t core = (get_absolute_logical_y() - Y0) * GW + (get_absolute_logical_x() - X0); + const uint32_t w = core * SLOTS + SLOT; + if (w >= NUM_WORKERS) { + return; + } + + const auto xa = TensorAccessor(x_args, x_addr, IN_PAGE_BYTES); + const auto oa = TensorAccessor(off_args, off_addr, OFF_PAGE_BYTES); + const auto ya = TensorAccessor(out_args, out_addr, OUT_ROW_BYTES); + + const uint32_t data_l1 = align64(get_write_ptr(CB_DATA)); + const uint32_t off_l1 = align64(get_write_ptr(CB_OFF)); + const uint32_t probe_l1 = align64(get_write_ptr(CB_PROBE)); + const uint32_t out_l1 = align64(get_write_ptr(CB_OUT)); + int32_t* acc = reinterpret_cast(get_write_ptr(CB_ACC)); + + // ---- offsets: single 64-byte probes (binary search) and a streamed window (the worker's segments) + auto probe = [&](uint32_t s) -> uint32_t { + const uint32_t byte = s * 4u; + noc_async_read(oa.get_noc_addr(0, byte & ~63u), probe_l1, 64); + noc_async_read_barrier(); + COMPILER_BARRIER(); + const uint32_t v = reinterpret_cast(probe_l1)[(byte & 63u) >> 2]; + return v < NUM_ROWS ? v : NUM_ROWS; + }; + uint32_t win0 = 0xffffffffu; // first offset index held in the window (multiple of 16) + auto off_at = [&](uint32_t s) -> uint32_t { + if (win0 == 0xffffffffu || s < win0 || s >= win0 + OFF_CHUNK) { + win0 = s & ~15u; + const uint32_t start = win0 * 4u; + uint32_t bytes = OFF_CHUNK * 4u; + if (start + bytes > OFF_ALIGNED_BYTES) { + bytes = OFF_ALIGNED_BYTES - start; + } + noc_async_read(oa.get_noc_addr(0, start), off_l1, bytes); + noc_async_read_barrier(); + COMPILER_BARRIER(); + } + const uint32_t v = reinterpret_cast(off_l1)[s - win0]; + return v < NUM_ROWS ? v : NUM_ROWS; + }; + + // ---- this worker's segments [s_lo, s_hi): costs off[s] + s in [w T / W, (w + 1) T / W) + const uint32_t total = probe(NUM_SEG) + NUM_SEG; + auto search = [&](uint32_t target) -> uint32_t { // smallest s in [0, NUM_SEG] with off[s] + s >= target + uint32_t lo = 0, hi = NUM_SEG; + while (lo < hi) { + const uint32_t mid = (lo + hi) >> 1; + if (probe(mid) + mid >= target) { + hi = mid; + } else { + lo = mid + 1; + } + } + return lo; + }; + const uint32_t t0 = static_cast((static_cast(w) * total) / NUM_WORKERS); + const uint32_t t1 = static_cast((static_cast(w + 1) * total) / NUM_WORKERS); + const uint32_t s_lo = (w == 0) ? 0 : search(t0); + const uint32_t s_hi = (w + 1 == NUM_WORKERS) ? NUM_SEG : search(t1); + if (s_lo >= s_hi) { + return; + } + + // ---- input rows: units of 32 rows, two L1 slots, the next unit prefetched while the current one is reduced + uint32_t row_last = off_at(s_lo); // largest row this worker reads (+1), monotone over its segments + { + const uint32_t r_end = off_at(s_hi); + if (r_end > row_last) { + row_last = r_end; + } + } + const uint32_t unit_last = row_last > 0 ? (row_last - 1) / UNIT_ROWS : 0; + auto issue_unit = [&](uint32_t u, uint32_t slot) { + const uint32_t dst = data_l1 + slot * UNIT_BYTES; + const uint32_t first_page = IN_TILE ? u * CT : u * UNIT_ROWS; + for (uint32_t p = 0; p < PAGES_PER_UNIT; ++p) { + noc_async_read(xa.get_noc_addr(first_page + p), dst + p * IN_PAGE_BYTES, IN_PAGE_BYTES); + } + }; + uint32_t cur_unit = 0xffffffffu, cur_slot = 0; + bool next_ready = false; // the unit cur_unit + 1 is in flight into slot cur_slot ^ 1 + auto ensure_unit = [&](uint32_t u) { + if (u == cur_unit) { + return; + } + if (next_ready && u == cur_unit + 1) { + noc_async_read_barrier(); + cur_slot ^= 1u; + } else { + noc_async_read_barrier(); // drain a stale prefetch before its slot is reused + issue_unit(u, cur_slot); + noc_async_read_barrier(); + } + COMPILER_BARRIER(); + cur_unit = u; + next_ready = false; + if (u + 1 <= unit_last && u + 1 < NUM_UNITS) { + issue_unit(u + 1, cur_slot ^ 1u); + next_ready = true; + } + }; + + auto reset_acc = [&]() { + for (uint32_t c = 0; c < C; ++c) { + acc[c] = static_cast(0x80000000u); + } + }; + auto accumulate_row = [&](uint32_t r) { // acc = max(acc, x[r, :]) on keys + const uint32_t i = r - cur_unit * UNIT_ROWS; + const uint32_t base = data_l1 + cur_slot * UNIT_BYTES; + if constexpr (IN_TILE) { + const uint32_t face_row = (i >> 4) << 1; // faces 0/1 (rows 0-15) or 2/3 (rows 16-31) + for (uint32_t tc = 0; tc < CT; ++tc) { + for (uint32_t h = 0; h < 2; ++h) { + const uint32_t e0 = tc * TILE_ELEMS + (face_row + h) * 256 + (i & 15u) * 16; + int32_t* a = acc + tc * 32 + h * 16; + if constexpr (IN_BYTES == 4) { + volatile tt_l1_ptr uint32_t* src = reinterpret_cast(base) + e0; + for (uint32_t k = 0; k < 16; ++k) { + const int32_t key = to_key(src[k]); + a[k] = key > a[k] ? key : a[k]; + } + } else { + volatile tt_l1_ptr uint16_t* src = reinterpret_cast(base) + e0; + for (uint32_t k = 0; k < 16; ++k) { + const int32_t key = to_key(static_cast(src[k]) << 16); + a[k] = key > a[k] ? key : a[k]; + } + } + } + } + } else { + if constexpr (IN_BYTES == 4) { + volatile tt_l1_ptr uint32_t* src = reinterpret_cast(base) + i * C; + for (uint32_t c = 0; c < C; ++c) { + const int32_t key = to_key(src[c]); + acc[c] = key > acc[c] ? key : acc[c]; + } + } else { + volatile tt_l1_ptr uint16_t* src = reinterpret_cast(base) + i * C; + for (uint32_t c = 0; c < C; ++c) { + const int32_t key = to_key(static_cast(src[c]) << 16); + acc[c] = key > acc[c] ? key : acc[c]; + } + } + } + }; + + // ---- output rows: a ring of OUT_RING staging rows, reused after noc_async_writes_flushed + uint32_t emitted = 0; + auto emit = [&](uint32_t s, bool empty) { + const uint32_t ring = emitted % OUT_RING; + if (ring == 0 && emitted > 0) { + noc_async_writes_flushed(); + } + const uint32_t dst = out_l1 + ring * OUT_ROW_BYTES; + if constexpr (OUT_BYTES == 4) { + volatile tt_l1_ptr uint32_t* o = reinterpret_cast(dst); + for (uint32_t c = 0; c < C; ++c) { + o[c] = empty ? FILL_BITS : from_key(acc[c]); + } + } else { + volatile tt_l1_ptr uint16_t* o = reinterpret_cast(dst); + const uint16_t fill = to_bf16_rne(FILL_BITS); + for (uint32_t c = 0; c < C; ++c) { + o[c] = empty ? fill : to_bf16_rne(from_key(acc[c])); + } + } + COMPILER_BARRIER(); + noc_async_write(dst, ya.get_noc_addr(s), OUT_ROW_BYTES); + ++emitted; + }; + + // ---- the reduction + uint32_t seg_begin = off_at(s_lo); + for (uint32_t s = s_lo; s < s_hi; ++s) { + uint32_t seg_end = off_at(s + 1); + if (seg_end < seg_begin) { + seg_end = seg_begin; // non-decreasing offsets (a corrupt table reads nothing extra) + } + if (seg_end == seg_begin) { + emit(s, true); + } else { + reset_acc(); + for (uint32_t r = seg_begin; r < seg_end; ++r) { + ensure_unit(r / UNIT_ROWS); + accumulate_row(r); + } + emit(s, false); + } + seg_begin = seg_end; + } + noc_async_read_barrier(); // a prefetch past the last row may still be in flight + noc_async_write_barrier(); +} diff --git a/code/tt_diffusion_planner/ttaw/ops/segment.py b/code/tt_diffusion_planner/ttaw/ops/segment.py new file mode 100644 index 0000000000000000000000000000000000000000..3cded44cf024976b18f83d5b5e9d106e0d93a072 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/ops/segment.py @@ -0,0 +1,447 @@ +# SPDX-License-Identifier: Apache-2.0 +"""K1 ``segment_reduce``: per-segment reductions over rows sorted by segment (PLAN.md section 1.3 row K1; owned by the +FRNet port), with the exact stock-op fallback (the log-step segmented max) and their host twins. + +A *segmented* tensor holds rows sorted by segment id: segment ``s`` is the contiguous row range +``[offsets[s], offsets[s + 1])`` (CSR offsets, non-decreasing; rows past ``offsets[-1]`` are padding that nothing +reads). FRNet's six ``ScatterElements(max)`` sites (points sorted by frustum cell, PLAN.md 2.8), PTv3 pooling and +BEVPool are of this form once the host sorts the rows. + +- **K1 on the device** (:class:`SegmentReduce`, kernel ``kernels/segment_reduce_dm.cpp``): ``out[s] = + max(x[offsets[s]:offsets[s + 1]])`` for ``s < num_segments``, empty segments -> ``empty_value`` (FRNet: the zero + row of the frustum2pixel gather table comes out of the kernel itself). Input ``[1, 1, R, C]`` TILE or ROW_MAJOR, + bf16 or fp32 (``R % 32 == 0``, ``C % 32 == 0``); offsets uint32 ``[1, K]`` ROW_MAJOR (``K >= num_segments + 1``, + an RT-dev tensor: rewritten per frame, no recapture); output ``[1, 1, num_segments, C]`` ROW_MAJOR, bf16 (round to + nearest even) or fp32 -- the layout ``ttnn.embedding`` takes as a table. **Exact**: the maximum is taken on + integer keys of the float bit patterns (signed order == float order, -0 < +0), so the result is bit-exact for any + input dtype, and fp32 -> bf16 output equals rounding first (a monotonic rounding commutes with max). One program, + common runtime args only (probe P3 rule), deterministic, no atomics, traceable. Work split: 2 data-movement + RISC-Vs per core, each a contiguous run of segments balanced on rows + 1 per segment, found on the device by binary + search over the offsets (no per-frame host split, nothing grid-dependent in the host tables). +- **max only on the device for now.** ``sum`` / ``mean`` need float adds, which the data-movement RISC-Vs do not have + (no F extension); they are host-oracle only here (:func:`segment_reduce_numpy`) and come with their consumers + (PTv3 / BEVPool, optimization phase). +- **Stock-op fallback** (:func:`logstep_segment_max`, PLAN.md 1.3): ``R`` Hillis-Steele rounds ``x = max(x, x[idx_r])`` + with host shift tables ``idx_r[i] = i - 2^r`` when row ``i - 2^r`` is in the same segment, else ``i`` (max is + idempotent), then one gather of each segment's last row: exact for segments of up to ``2^R`` rows, ``3R + 1`` + programs (untilize, ``ttnn.embedding``, ``ttnn.maximum`` per round). ``ttnn.embedding`` is bf16-only, so the + fallback is bf16 in and out. ``R`` is COMPILE (it sets the op count): bucket it per frame (:func:`segment_rounds`). + +Host side (numpy only): :func:`segment_offsets` (CSR offsets from sorted segment ids), :func:`segment_rounds`, +:func:`segment_shift_tables`, :func:`segment_last_rows`, oracles :func:`segment_reduce_numpy` and +:func:`logstep_segment_max_numpy`, and :func:`worker_segments` (the kernel's work split, for tests and diagnostics). +numpy only at import; ttnn is imported inside the device functions. +""" +from __future__ import annotations + +from dataclasses import dataclass +from typing import Any, Dict, List, Optional, Sequence, Tuple + +import numpy as np + +from . import kernel_path + +__all__ = [ + "OPS", + "DEVICE_OPS", + "segment_offsets", + "check_offsets", + "segment_rounds", + "segment_shift_tables", + "segment_last_rows", + "segment_reduce_numpy", + "logstep_segment_max_numpy", + "worker_segments", + "SegmentReduceSpec", + "SegmentReduce", + "segment_reduce", + "logstep_segment_max", + "KERNEL", +] + +OPS = ("max", "min", "sum", "mean") +DEVICE_OPS = ("max",) +KERNEL = "segment_reduce_dm.cpp" +UNIT_ROWS = 32 # rows per streamed unit (one tile-row): the input rows must be a multiple of it +_OFF_CHUNK = 256 # offsets per L1 window (multiple of 16) +_OUT_RING = 8 # output staging rows in flight per worker + + +# ------------------------------------------------------------------------------------------------ host tables + +def segment_offsets(seg_id: Any, num_segments: int, *, num_rows: Optional[int] = None) -> np.ndarray: + """CSR offsets ``uint32 [num_segments + 1]`` of rows sorted by segment id: segment ``s`` = rows + ``[off[s], off[s + 1])``. ``seg_id`` (``int [R]``, non-decreasing) may contain ids ``>= num_segments`` (padding + rows, e.g. FRNet's trailing sentinel segment): they must come last and belong to no segment. ``num_rows`` + (default ``len(seg_id)``) bounds the offsets.""" + ids = np.asarray(seg_id).reshape(-1).astype(np.int64) + if ids.size and (np.any(np.diff(ids) < 0) or ids[0] < 0): + raise ValueError("segment ids must be non-negative and sorted (non-decreasing)") + counts = np.bincount(ids[ids < num_segments], minlength=num_segments).astype(np.int64) + off = np.zeros(num_segments + 1, np.int64) + np.cumsum(counts, out=off[1:]) + rows = len(ids) if num_rows is None else int(num_rows) + if off[-1] > rows: + raise ValueError(f"segments hold {off[-1]} rows, more than num_rows={rows}") + return off.astype(np.uint32) + + +def check_offsets(offsets: Any, *, num_segments: int, num_rows: int, length: Optional[int] = None) -> np.ndarray: + """Validate CSR offsets for the device and return them as a contiguous ``uint32 [1, length]`` row (``length`` + defaults to ``num_segments + 1``; extra entries repeat the last offset). Raises ``ValueError`` unless they start + at 0, never decrease and end at most at ``num_rows`` (the kernel clamps anyway; a bad table is a host bug).""" + a = np.asarray(offsets).reshape(-1) + if a.dtype.kind not in "iu": + raise ValueError(f"offsets must be integral, got dtype {a.dtype}") + a = a.astype(np.int64) + if a.size < num_segments + 1: + raise ValueError(f"{a.size} offsets for {num_segments} segments (need num_segments + 1)") + a = a[: num_segments + 1] + if a[0] != 0 or np.any(np.diff(a) < 0) or a[-1] > num_rows: + raise ValueError(f"offsets must start at 0, be non-decreasing and end <= num_rows={num_rows}") + n = num_segments + 1 if length is None else int(length) + if n < num_segments + 1: + raise ValueError(f"length {n} < num_segments + 1") + out = np.full((1, n), a[-1], np.uint32) + out[0, : a.size] = a + return out + + +def segment_rounds(max_segment_length: int) -> int: + """Log-step rounds that cover segments of up to ``max_segment_length`` rows: ``ceil(log2(L))`` (0 for L <= 1).""" + L = int(max_segment_length) + return int(np.ceil(np.log2(L))) if L > 1 else 0 + + +def segment_shift_tables(seg_id: Any, rounds: int) -> List[np.ndarray]: + """The gather tables of the log-step segmented max: round ``r`` uses ``idx[i] = i - 2^r`` if row ``i - 2^r`` has + the same segment id, else ``i`` -> ``rounds`` x ``uint32 [R]``. After ``rounds`` rounds the last row of every + segment of at most ``2^rounds`` rows holds the segment max (padding rows only ever see padding rows).""" + seg = np.asarray(seg_id).reshape(-1) + i = np.arange(len(seg), dtype=np.int64) + out = [] + for r in range(int(rounds)): + d = 1 << r + same = np.zeros(len(seg), bool) + same[d:] = seg[d:] == seg[:-d] + out.append(np.where(same, i - d, i).astype(np.uint32)) + return out + + +def segment_last_rows(offsets: Any, *, empty_row: int) -> np.ndarray: + """``uint32 [S]``: the last row of each segment of ``offsets`` (``S + 1`` entries), ``empty_row`` for an empty + segment. For :func:`logstep_segment_max` pass the number of input rows: the fallback appends one zero row there + (padding rows themselves are reduced too, so they do not stay zero).""" + off = np.asarray(offsets).reshape(-1).astype(np.int64) + last = off[1:] - 1 + return np.where(off[1:] > off[:-1], last, int(empty_row)).astype(np.uint32) + + +# ------------------------------------------------------------------------------------------------- oracles + +def _float_keys(a: np.ndarray) -> np.ndarray: + """float32 -> int32 keys whose signed order is the float order with -0 < +0 (an involution on the bits).""" + b = np.ascontiguousarray(a, np.float32).view(np.int32) + return b ^ ((b >> 31) & np.int32(0x7FFFFFFF)) + + +def _from_keys(k: np.ndarray) -> np.ndarray: + return (k ^ ((k >> 31) & np.int32(0x7FFFFFFF))).astype(np.int32).view(np.float32) + + +def segment_reduce_numpy(x: Any, offsets: Any, num_segments: Optional[int] = None, *, op: str = "max", + empty_value: float = 0.0) -> np.ndarray: + """Host oracle: ``out[s] = op(x[off[s]:off[s + 1]], axis=0)`` (``x`` ``[R, C]`` or ``[..., R, C]``), empty + segments -> ``empty_value``. ``max`` / ``min`` keep the input dtype (exact) and, for float32 / bf16-valued + inputs, order the floats as the K1 kernel does (signed zeros: ``-0 < +0``; numpy's ``maximum`` would keep + whichever zero comes first); ``sum`` / ``mean`` accumulate in float64 and return float32.""" + if op not in OPS: + raise ValueError(f"op {op!r}: one of {OPS}") + a = np.asarray(x) + a = a.reshape(-1, a.shape[-1]) + off = np.asarray(offsets).reshape(-1).astype(np.int64) + S = len(off) - 1 if num_segments is None else int(num_segments) + off = off[: S + 1] + out_dtype = a.dtype if op in ("max", "min") else np.float32 + out = np.full((S, a.shape[1]), empty_value, dtype=out_dtype) + nonempty = np.flatnonzero(off[1:] > off[:-1]) + if nonempty.size: + starts = off[:-1][nonempty] + if op in ("max", "min") and a.dtype == np.float32: + keys = _float_keys(a[: off[-1]]) + red = _from_keys((np.maximum if op == "max" else np.minimum).reduceat(keys, starts, axis=0)) + elif op in ("max", "min"): + red = (np.maximum if op == "max" else np.minimum).reduceat(a[: off[-1]], starts, axis=0) + else: + red = np.add.reduceat(a[: off[-1]].astype(np.float64), starts, axis=0) + if op == "mean": + red = red / (off[1:][nonempty] - starts)[:, None] + out[nonempty] = red.astype(out_dtype) + return out + + +def logstep_segment_max_numpy(x: Any, shift_tables: Sequence[Any], last_rows: Any) -> np.ndarray: + """Host twin of :func:`logstep_segment_max` (any dtype; exact): the rounds, then a gather of ``last_rows`` from + the result with one zero row appended (index ``R``: the empty segments).""" + a = np.asarray(x) + a = a.reshape(-1, a.shape[-1]) + for idx in shift_tables: + a = np.maximum(a, a[np.asarray(idx).reshape(-1).astype(np.int64)]) + a = np.concatenate([a, np.zeros((1, a.shape[1]), a.dtype)], axis=0) + return a[np.asarray(last_rows).reshape(-1).astype(np.int64)] + + +def worker_segments(offsets: Any, num_segments: int, num_workers: int, *, num_rows: Optional[int] = None + ) -> List[Tuple[int, int]]: + """The kernel's work split (host twin of ``segment_reduce_dm.cpp``): worker ``w`` owns the segments whose cost + ``off[s] + s`` lies in ``[w T / W, (w + 1) T / W)`` with ``T = off[S] + S`` -> ``[(s_lo, s_hi)]`` per worker. + Offsets are clamped to ``num_rows`` as the kernel clamps them.""" + off = np.asarray(offsets).reshape(-1).astype(np.int64)[: num_segments + 1] + if num_rows is not None: + off = np.minimum(off, int(num_rows)) + cost = off + np.arange(num_segments + 1) + total = int(cost[-1]) + out = [] + for w in range(num_workers): + t0 = (w * total) // num_workers + t1 = ((w + 1) * total) // num_workers + lo = 0 if w == 0 else int(np.searchsorted(cost, t0, side="left")) + hi = num_segments if w + 1 == num_workers else int(np.searchsorted(cost, t1, side="left")) + out.append((lo, max(lo, hi))) + return out + + +# ------------------------------------------------------------------------------------------------- device K1 + +def _f32_bits(value: float) -> int: + return int(np.array([value], np.float32).view(np.uint32)[0]) + + +@dataclass(frozen=True) +class SegmentReduceSpec: + """The static shape of one K1 call site (COMPILE): ``num_rows`` x ``channels`` input, ``num_segments`` outputs, + dtypes / layout, the empty value, and how many workers (data-movement RISC-Vs) per core run it.""" + + num_rows: int + channels: int + num_segments: int + input_dtype: str = "float32" # "float32" | "bfloat16" + input_layout: str = "tile" # "tile" | "row_major" + output_dtype: str = "float32" # "float32" | "bfloat16" (RNE) + op: str = "max" + empty_value: float = 0.0 + workers_per_core: int = 2 + + def __post_init__(self) -> None: + if self.op not in DEVICE_OPS: + raise NotImplementedError(f"K1 runs {DEVICE_OPS} on the device; {self.op!r} is host-oracle only " + "(segment_reduce_numpy): sum / mean need float adds the data-movement RISC-Vs " + "do not have") + if self.num_rows <= 0 or self.num_rows % UNIT_ROWS: + raise ValueError(f"num_rows must be a positive multiple of {UNIT_ROWS}, got {self.num_rows}") + if self.channels <= 0 or self.channels % 32: + raise ValueError(f"channels must be a positive multiple of 32, got {self.channels}") + if self.num_segments <= 0: + raise ValueError("num_segments must be positive") + for name, value in (("input_dtype", self.input_dtype), ("output_dtype", self.output_dtype)): + if value not in ("float32", "bfloat16"): + raise ValueError(f"{name} {value!r}: 'float32' or 'bfloat16'") + if self.input_layout not in ("tile", "row_major"): + raise ValueError(f"input_layout {self.input_layout!r}: 'tile' or 'row_major'") + if self.workers_per_core not in (1, 2): + raise ValueError("workers_per_core must be 1 or 2") + if not np.isfinite(self.empty_value): + raise ValueError("empty_value must be finite") + + @property + def in_bytes(self) -> int: + return 4 if self.input_dtype == "float32" else 2 + + @property + def out_bytes(self) -> int: + return 4 if self.output_dtype == "float32" else 2 + + @property + def in_page_bytes(self) -> int: + return (1024 if self.input_layout == "tile" else self.channels) * self.in_bytes + + @property + def unit_bytes(self) -> int: + pages = self.channels // 32 if self.input_layout == "tile" else UNIT_ROWS + return pages * self.in_page_bytes + + @property + def out_row_bytes(self) -> int: + return self.channels * self.out_bytes + + def l1_bytes_per_worker(self) -> Dict[str, int]: + """The L1 scratch (CB) sizes of one worker: 2 data units, offsets window, probe block, accumulator, output + ring (each + 64 B so the kernel can align to the 64-byte DRAM read alignment).""" + return {"data": 2 * self.unit_bytes + 64, "offsets": _OFF_CHUNK * 4 + 64, "probe": 128, + "acc": self.channels * 4, "out": _OUT_RING * self.out_row_bytes + 64} + + def check_l1(self, budget: int = 600 * 1024) -> None: + """Refuse configurations whose per-core scratch exceeds ``budget`` (default 600 KiB of ~1.4 MiB).""" + per_core = self.workers_per_core * sum(self.l1_bytes_per_worker().values()) + if per_core > budget: + raise ValueError(f"K1 scratch {per_core} B per core exceeds {budget} B") + + +class SegmentReduce: + """K1 for one call site (static :class:`SegmentReduceSpec`). ``__call__(x, offsets, *, output=None)`` runs one + ``ttnn.generic_op`` (eagerly or inside a trace capture) and returns the ``[1, 1, num_segments, C]`` ROW_MAJOR + DRAM output (allocated per call unless ``output`` is given: a persistent buffer of that spec). + + ``x``: ``[1, 1, num_rows, C]`` interleaved DRAM tensor of the spec's dtype / layout, rows sorted by segment; + ``offsets``: uint32 ROW_MAJOR ``[1, K]`` DRAM (``K >= num_segments + 1``; :func:`check_offsets`), e.g. a + ``TraceRunner`` input rewritten per frame. The core grid is read from the device (never hard-coded); the split + across cores happens on the device, so ETH (12x10) and WORKER (11x10) runs give identical outputs.""" + + def __init__(self, spec: SegmentReduceSpec, *, name: str = "segment_reduce", grid: Optional[Tuple[int, int]] = None): + self.spec = spec + self.name = name + self.grid = grid # (x, y) cores; None = the device's compute grid + spec.check_l1() + + # ---- descriptor + def _grid(self, device: Any) -> Tuple[int, int]: + if self.grid is not None: + return int(self.grid[0]), int(self.grid[1]) + g = device.compute_with_storage_grid_size() + return int(g.x), int(g.y) + + def program(self, x: Any, offsets: Any, out: Any): + """The ``ttnn.ProgramDescriptor`` of one call (buffer addresses in common runtime args).""" + import ttnn + + spec = self.spec + gx, gy = self._grid(x.device()) + rect = ttnn.CoreRange(ttnn.CoreCoord(0, 0), ttnn.CoreCoord(gx - 1, gy - 1)) + crs = ttnn.CoreRangeSet([rect]) + slots = spec.workers_per_core + workers = gx * gy * slots + off_len = int(offsets.shape[-1]) + sizes = spec.l1_bytes_per_worker() + cbs, kernels = [], [] + configs = [ttnn.WriterConfigDescriptor(), ttnn.ReaderConfigDescriptor()][:slots] + accessors = (list(ttnn.TensorAccessorArgs(x).get_compile_time_args()) + + list(ttnn.TensorAccessorArgs(offsets).get_compile_time_args()) + + list(ttnn.TensorAccessorArgs(out).get_compile_time_args())) + for slot, config in enumerate(configs): + ids = [5 * slot + k for k in range(5)] # data, offsets, probe, acc, out + for cb_id, size in zip(ids, (sizes["data"], sizes["offsets"], sizes["probe"], sizes["acc"], + sizes["out"])): + cbs.append(ttnn.CBDescriptor(total_size=size, core_ranges=crs, format_descriptors=[ + ttnn.CBFormatDescriptor(buffer_index=cb_id, data_format=ttnn.uint32, page_size=size)])) + ct = [1 if spec.input_layout == "tile" else 0, spec.in_bytes, spec.out_bytes, spec.channels, + spec.num_rows, spec.num_segments, workers, slots, slot, 0, 0, gx, *ids, _OUT_RING, _OFF_CHUNK, + _f32_bits(spec.empty_value), off_len * 4] + accessors + kernels.append(ttnn.KernelDescriptor( + kernel_source=kernel_path(KERNEL), source_type=ttnn.KernelDescriptor.SourceType.FILE_PATH, + core_ranges=crs, compile_time_args=ct, defines=[], runtime_args=ttnn.RuntimeArgs(), + common_runtime_args=[x.buffer_address(), offsets.buffer_address(), out.buffer_address()], + config=config)) + return ttnn.ProgramDescriptor(kernels=kernels, semaphores=[], cbs=cbs) + + # ---- checks + def _check_inputs(self, x: Any, offsets: Any) -> None: + import ttnn + + spec = self.spec + shape = tuple(int(s) for s in x.shape) + if shape != (1, 1, spec.num_rows, spec.channels): + raise ValueError(f"{self.name}: x is {shape}, expected (1, 1, {spec.num_rows}, {spec.channels})") + want_dtype = ttnn.float32 if spec.input_dtype == "float32" else ttnn.bfloat16 + if x.dtype != want_dtype: + raise ValueError(f"{self.name}: x dtype {x.dtype}, expected {spec.input_dtype}") + want_layout = ttnn.TILE_LAYOUT if spec.input_layout == "tile" else ttnn.ROW_MAJOR_LAYOUT + if x.layout != want_layout: + raise ValueError(f"{self.name}: x layout {x.layout}, expected {spec.input_layout}") + oshape = tuple(int(s) for s in offsets.shape) + if len(oshape) != 2 or oshape[0] != 1 or oshape[1] < spec.num_segments + 1: + raise ValueError(f"{self.name}: offsets must be [1, K >= {spec.num_segments + 1}], got {oshape}") + if offsets.dtype != ttnn.uint32 or offsets.layout != ttnn.ROW_MAJOR_LAYOUT: + raise ValueError(f"{self.name}: offsets must be uint32 ROW_MAJOR") + for t, what in ((x, "x"), (offsets, "offsets")): + mc = t.memory_config() + if getattr(mc, "buffer_type", ttnn.BufferType.DRAM) != ttnn.BufferType.DRAM or \ + getattr(mc, "memory_layout", ttnn.TensorMemoryLayout.INTERLEAVED) != \ + ttnn.TensorMemoryLayout.INTERLEAVED: + raise ValueError(f"{self.name}: {what} must be DRAM interleaved") + + def output_spec(self): + """``(shape, ttnn dtype, layout)`` of the output.""" + import ttnn + + dt = ttnn.float32 if self.spec.output_dtype == "float32" else ttnn.bfloat16 + return (1, 1, self.spec.num_segments, self.spec.channels), dt, ttnn.ROW_MAJOR_LAYOUT + + def allocate_output(self, device: Any): + import ttnn + + shape, dt, layout = self.output_spec() + return ttnn.allocate_tensor_on_device(ttnn.Shape(list(shape)), dt, layout, device, ttnn.DRAM_MEMORY_CONFIG) + + def __call__(self, x: Any, offsets: Any, *, output: Any = None): + import ttnn + + self._check_inputs(x, offsets) + out = self.allocate_output(x.device()) if output is None else output + shape, dt, layout = self.output_spec() + if tuple(int(s) for s in out.shape) != shape or out.dtype != dt or out.layout != layout: + raise ValueError(f"{self.name}: output must be {shape} {dt} ROW_MAJOR") + return ttnn.generic_op([x, offsets, out], self.program(x, offsets, out)) + + def describe(self) -> Dict[str, Any]: + s = self.spec + return {"name": self.name, "kernel": KERNEL, "op": s.op, "num_rows": s.num_rows, "channels": s.channels, + "num_segments": s.num_segments, "input": f"{s.input_dtype}/{s.input_layout}", + "output": f"{s.output_dtype}/row_major", "empty_value": s.empty_value, + "workers_per_core": s.workers_per_core, "grid": list(self.grid) if self.grid else "device", + "l1_bytes_per_worker": s.l1_bytes_per_worker()} + + +def segment_reduce(x: Any, offsets: Any, *, num_segments: int, op: str = "max", output_dtype: Optional[str] = None, + empty_value: float = 0.0, output: Any = None): + """One-shot K1 (builds a :class:`SegmentReduce` from the tensors; ports build it once per call site).""" + import ttnn + + in_dtype = "float32" if x.dtype == ttnn.float32 else "bfloat16" + spec = SegmentReduceSpec(num_rows=int(x.shape[-2]), channels=int(x.shape[-1]), num_segments=int(num_segments), + input_dtype=in_dtype, input_layout="tile" if x.layout == ttnn.TILE_LAYOUT else + "row_major", output_dtype=output_dtype or in_dtype, op=op, empty_value=empty_value) + return SegmentReduce(spec)(x, offsets, output=output) + + +# -------------------------------------------------------------------------------------- stock-op fallback + +def logstep_segment_max(x: Any, shift_tables: Sequence[Any], last_rows: Any, *, layout: str = "tile"): + """The exact log-step segmented max with stock ops (PLAN.md 1.3; the K1 fallback): ``x`` bf16 ``[1, 1, R, C]`` + (TILE or ROW_MAJOR, ``C % 32 == 0``), ``shift_tables`` = ``rounds`` uint32 ``[1, R]`` device tensors + (:func:`segment_shift_tables`), ``last_rows`` uint32 ``[1, S]`` (:func:`segment_last_rows` with ``empty_row=R``) + -> bf16 ``[1, 1, S, C]`` (``layout``). Per round: untilize, ``ttnn.embedding`` gather, ``ttnn.maximum`` (3 + programs); then one PADDED gather from the result with a zero row appended at ``R`` (empty segments -> 0). + bf16 only (``ttnn.embedding`` tables are bf16); exact for segments of up to ``2^rounds`` rows.""" + import ttnn + + from .gather import gather_rows, with_zero_rows + + if x.dtype != ttnn.bfloat16: + raise ValueError("logstep_segment_max is bf16-only (ttnn.embedding); use K1 (SegmentReduce) for fp32") + cur = x if x.layout == ttnn.TILE_LAYOUT else ttnn.to_layout(x, ttnn.TILE_LAYOUT) + for idx in shift_tables: + table = ttnn.to_layout(cur, ttnn.ROW_MAJOR_LAYOUT) + shifted = gather_rows(idx, table, layout="tile") + nxt = ttnn.maximum(cur, shifted) + for t in (table, shifted): + ttnn.deallocate(t, False) + if cur is not x: + ttnn.deallocate(cur, False) + cur = nxt + rows = int(cur.shape[-2]) + table = with_zero_rows(cur) # [1, 1, R + 1, C] ROW_MAJOR, row R = 0 + out = gather_rows(last_rows, table, sentinel=rows, layout=layout) + ttnn.deallocate(table, False) + if cur is not x: + ttnn.deallocate(cur, False) + return out diff --git a/code/tt_diffusion_planner/ttaw/ops/topk.py b/code/tt_diffusion_planner/ttaw/ops/topk.py new file mode 100644 index 0000000000000000000000000000000000000000..689dfe4cf7f1103a07b8dd95575abde0f817be31 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/ops/topk.py @@ -0,0 +1,272 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C21: top-k query selection on the device (PLAN.md section 1.2 row C21; probe P8, common/PROBES.md). + +The query-based LiDAR detectors (TransFusion, BEVFusion, PTv3) pick the ``K`` best (class, cell) proposals of a +local-max heatmap ``[C, H*W]``. On the p150 that is ``ttnn.experimental.topk_large_indices`` (Blackhole only): one +bf16 ROW_MAJOR row, ``k`` a multiple of 16 in [16, 2048], uint32 indices in descending value order, ~4.3 ns per +element on one core (0.71 / 0.80 / 1.98 ms over 162k / 184k / 459k values, P8). What the helpers here guarantee: + +- **The row is class-major** (:func:`class_major_row`): index ``c * cells + p`` for class ``c`` and cell ``p`` -- + the ONNX ``TopK`` order of a ``[1, C, H*W]`` reshape, so reference (teacher-forcing) indices need no conversion. + The NHWC heatmap ``[1, 1, cells, C]`` (TILE) is transposed (one program), reshaped to one row while still tiled, + and untilized: no ROW_MAJOR reshape of wide pages, whose L1 staging (2 x the destination page, API.md 13) would + not fit PT's 0.9 MB row (TF's is 0.37 MB). +- **Exact decode** (:func:`decode_class_major`): the uint32 indices are cast to fp32 (exact below 2**24) and decoded + with comparisons only, ``cls = sum_c (idx >= c * cells)``, ``pos = idx - cls * cells`` -- no division, so no + device rounding can move a boundary -- clamped into range (a 0xFFFFFFFF sentinel can never become an + out-of-range gather index) and cast back to uint32 ROW_MAJOR ``[1, K]`` rows (``ttnn.embedding`` indices). +- **Ties are unspecified** (P8: the device does not break ties by lowest index), so every gate after the selection + is set-based (:func:`selection_metrics`) or teacher-forced; the host oracle :func:`topk_numpy` has ONNX semantics + (stable descending: ties to the lower index). +- **Never ``-inf`` masks**: surviving ``-inf`` values come back as index 0xFFFFFFFF (P8); mask with + :data:`MASK_VALUE` (-1e30, finite in bf16) instead. +- ``k`` is padded to a multiple of 16 (:func:`padded_k`; 500 -> 512); keep the first ``K`` of the descending output. + +The bf16 ``ttnn.topk`` composite also returns the values (P8: +8-98 us, exact values): :func:`topk_with_values`. An +fp32 value re-gather costs ~1 us per index through ``ttnn.gather`` (P13); a two-stage top-k (row-split first stage) +is optimization work (PLAN.md 5.3). + +numpy only at import; ttnn is imported inside the device functions. +""" +from __future__ import annotations + +from typing import Any, Dict, List, Optional, Sequence, Tuple + +import numpy as np + +__all__ = ["K_MULTIPLE", "K_MIN", "K_MAX", "MASK_VALUE", "SENTINEL", "padded_k", "class_major_row", "topk_indices", + "decode_class_major", "topk_with_values", "TopKSelect", "topk_numpy", "class_major_numpy", + "decode_class_major_numpy", "selection_metrics", "near_tie_report"] + +K_MULTIPLE = 16 +K_MIN = 16 +K_MAX = 2048 +MASK_VALUE = -1.0e30 # masking value for scores (never -inf: P8 sentinel) +SENTINEL = 0xFFFFFFFF # index returned for a surviving -inf +FP32_EXACT = 1 << 24 # integers below are exact in fp32 (the decode's range) + + +def padded_k(k: int) -> int: + """``k`` rounded up to the multiple of 16 ``topk_large_indices`` takes (500 -> 512); 16 <= result <= 2048.""" + k = int(k) + if k <= 0: + raise ValueError(f"k must be positive, got {k}") + kp = max(K_MIN, -(-k // K_MULTIPLE) * K_MULTIPLE) + if kp > K_MAX: + raise ValueError(f"k={k} exceeds topk_large_indices' maximum {K_MAX}") + return kp + + +# ------------------------------------------------------------------------------------------------- device + +def class_major_row(heat_nhwc: Any, classes: int, cells: int, *, memory_config: Any = None): + """NHWC heatmap ``[1, 1, cells, classes]`` (TILE, bf16) -> the class-major bf16 ROW_MAJOR row + ``[1, classes * cells]`` of the top-k (transpose, TILE reshape, untilize: traceable).""" + import ttnn + + shape = tuple(int(s) for s in heat_nhwc.shape) + if shape != (1, 1, int(cells), int(classes)): + raise ValueError(f"heat must be [1, 1, {cells}, {classes}], got {shape}") + if heat_nhwc.dtype != ttnn.bfloat16: + raise ValueError(f"topk_large_indices takes bf16 scores, got {heat_nhwc.dtype}") + kw = {} if memory_config is None else {"memory_config": memory_config} + n = int(classes) * int(cells) + t = ttnn.transpose(heat_nhwc, 2, 3, **kw) # [1, 1, classes, cells] TILE + flat = ttnn.reshape(t, (1, 1, 1, n)) # the row while still tiled (TILE reshape) + ttnn.deallocate(t, False) + rm = ttnn.to_layout(flat, ttnn.ROW_MAJOR_LAYOUT) # one row, untilized + ttnn.deallocate(flat, False) + return ttnn.reshape(rm, (1, n)) + + +def topk_indices(row: Any, k: int): + """``topk_large_indices`` of a bf16 ROW_MAJOR ``[rows, n]`` tensor -> uint32 ROW_MAJOR ``[rows, padded_k(k)]``, + descending (ties unspecified). Refuses an ``n`` beyond fp32's exact integers (the decode's range).""" + import ttnn + + n = int(row.shape[-1]) + kp = padded_k(k) + if n < kp: + raise ValueError(f"the row has {n} values, fewer than k={kp}") + if n > FP32_EXACT: + raise ValueError(f"row of {n} values: indices above 2**24 are not exact in the fp32 decode") + if row.dtype != ttnn.bfloat16 or row.layout != ttnn.ROW_MAJOR_LAYOUT: + raise ValueError("topk_large_indices takes a bf16 ROW_MAJOR row") + return ttnn.experimental.topk_large_indices(row, k=kp) + + +def decode_class_major(indices: Any, cells: int, classes: int, *, clamp: bool = True) -> Tuple[Any, Any]: + """Class-major flat indices (uint32, ROW_MAJOR ``[1, K]`` or TILE) -> ``(pos, cls)`` uint32 ROW_MAJOR ``[1, K]`` + with ``idx = cls * cells + pos`` (module docstring: fp32 comparisons, exact; ``clamp`` keeps them in range).""" + import ttnn + + cells, classes = int(cells), int(classes) + if cells * classes > FP32_EXACT: + raise ValueError("cells * classes exceeds fp32's exact integers") + it = indices if indices.layout == ttnn.TILE_LAYOUT else ttnn.to_layout(indices, ttnn.TILE_LAYOUT) + f = ttnn.typecast(it, ttnn.float32) + if it is not indices: + ttnn.deallocate(it, False) + cls = None + for c in range(1, classes): + ge = ttnn.ge(f, float(c * cells)) + if cls is None: + cls = ge + else: + s = ttnn.add(cls, ge) + ttnn.deallocate(cls, False) + ttnn.deallocate(ge, False) + cls = s + if cls is None: # one class: cls = 0, pos = idx + pos = f + cls = ttnn.multiply(f, 0.0) + else: + off = ttnn.multiply(cls, float(cells)) + pos = ttnn.subtract(f, off) + ttnn.deallocate(off, False) + ttnn.deallocate(f, False) + if clamp: + p2 = ttnn.clamp(pos, 0.0, float(cells - 1)) + c2 = ttnn.clamp(cls, 0.0, float(classes - 1)) + ttnn.deallocate(pos, False) + ttnn.deallocate(cls, False) + pos, cls = p2, c2 + out = [] + for t in (pos, cls): + u = ttnn.typecast(t, ttnn.uint32) + ttnn.deallocate(t, False) + r = ttnn.to_layout(u, ttnn.ROW_MAJOR_LAYOUT) + ttnn.deallocate(u, False) + out.append(r) + return out[0], out[1] + + +def topk_with_values(scores_tile: Any, k: int): + """The bf16 ``ttnn.topk`` composite on a TILE ``[.., n]`` row: ``(values bf16, indices)`` of the top ``k`` + (P8: exact values, in-range indices even with ``-inf``; ties unspecified).""" + import ttnn + + return ttnn.topk(scores_tile, k=int(k), dim=-1, largest=True, sorted=True) + + +class TopKSelect: + """Top-``k`` of a ``[1, 1, cells, classes]`` bf16 TILE heatmap -> uint32 class-major indices ``[1, padded_k]`` + (descending) plus the decoded ``(pos, cls)`` rows: :func:`class_major_row`, :func:`topk_indices`, + :func:`decode_class_major`. ``keep`` (default ``k``) proposals are meaningful; the rows hold ``padded_k(k)``.""" + + def __init__(self, cells: int, classes: int, k: int, *, name: str = "topk"): + self.cells, self.classes, self.k = int(cells), int(classes), int(k) + self.kp = padded_k(self.k) + if self.cells * self.classes < self.kp: + raise ValueError(f"{name}: {self.cells * self.classes} scores for k={self.kp}") + self.name = name + + def indices(self, heat_nhwc: Any): + row = class_major_row(heat_nhwc, self.classes, self.cells) + idx = topk_indices(row, self.kp) + import ttnn + + ttnn.deallocate(row, False) + return idx + + def decode(self, indices: Any) -> Tuple[Any, Any]: + return decode_class_major(indices, self.cells, self.classes) + + def __call__(self, heat_nhwc: Any) -> Tuple[Any, Any, Any]: + """-> ``(indices, pos, cls)``, each uint32 ROW_MAJOR ``[1, padded_k]``.""" + idx = self.indices(heat_nhwc) + pos, cls = self.decode(idx) + return idx, pos, cls + + def describe(self) -> Dict[str, Any]: + return {"name": self.name, "cells": self.cells, "classes": self.classes, "k": self.k, "padded_k": self.kp, + "order": "class-major (idx = cls * cells + pos)", "ties": "unspecified (P8)"} + + +# --------------------------------------------------------------------------------------------------- host + +def topk_numpy(scores: Any, k: int) -> Tuple[np.ndarray, np.ndarray]: + """ONNX ``TopK(largest=1, sorted=1)`` of a 1-D score vector: ``(values, indices)``, values descending, ties in + ascending index order (a stable descending sort).""" + s = np.asarray(scores).reshape(-1) + order = np.argsort(-s.astype(np.float64), kind="stable")[: int(k)] + return s[order], order.astype(np.int64) + + +def class_major_numpy(heat_nhwc: Any) -> np.ndarray: + """NHWC rows ``[cells, classes]`` -> the class-major flat vector of :func:`class_major_row`.""" + h = np.asarray(heat_nhwc) + return np.ascontiguousarray(h.reshape(-1, h.shape[-1]).T).reshape(-1) + + +def decode_class_major_numpy(indices: Any, cells: int) -> Tuple[np.ndarray, np.ndarray]: + """Host decode: ``(pos, cls)`` int64 of class-major flat indices.""" + i = np.asarray(indices).astype(np.int64).reshape(-1) + return i % int(cells), i // int(cells) + + +def selection_metrics(device_indices: Any, reference_indices: Any, *, reference_scores: Optional[Any] = None, + flat_scores: Optional[Any] = None, strong: float = 0.1, k: Optional[int] = None + ) -> Dict[str, Any]: + """Set-based agreement of a device top-k with the reference's (PLAN.md 2.5 / 0.3 item 1). + + ``device_indices`` / ``reference_indices``: flat indices (the first ``k`` of each are compared, default the + reference length). ``reference_scores``: the reference values of ``reference_indices`` (same order) or + ``flat_scores``: the reference score of every flat index. Returns ``shared`` (count) / ``shared_frac``, and with + scores the strong proposals (reference score >= ``strong``): ``strong_total``, ``strong_kept``, + ``strong_overlap`` (fraction; 1.0 when there is none), ``strong_missed`` (indices).""" + d = np.asarray(device_indices).astype(np.int64).reshape(-1) + r = np.asarray(reference_indices).astype(np.int64).reshape(-1) + k = len(r) if k is None else int(k) + d, r = d[:k], r[:k] + ds = set(d.tolist()) + out: Dict[str, Any] = {"k": k, "shared": len(ds & set(r.tolist())), "unique": len(ds) == len(d)} + out["device_only"] = sorted(ds - set(r.tolist())) + out["shared_frac"] = out["shared"] / max(k, 1) + if reference_scores is not None or flat_scores is not None: + if reference_scores is not None: + vals = np.asarray(reference_scores, np.float64).reshape(-1)[:k] + strong_set = set(r[vals >= strong].tolist()) + else: + fs = np.asarray(flat_scores, np.float64).reshape(-1) + strong_set = set(np.flatnonzero(fs >= strong).tolist()) + missed = sorted(strong_set - ds) + out.update({"strong_threshold": float(strong), "strong_total": len(strong_set), + "strong_kept": len(strong_set) - len(missed), + "strong_overlap": 1.0 if not strong_set else (len(strong_set) - len(missed)) / len(strong_set), + "strong_missed": missed}) + return out + + +def near_tie_report(missed: Sequence[int], device_indices: Any, reference_heat: Any, height: int, width: int, + *, kernel: int = 3) -> List[Dict[str, Any]]: + """Why a reference proposal is missing from a device selection (diagnostics of :func:`selection_metrics`). + + ``reference_heat``: the reference's class-major sigmoid heat BEFORE the local max (``[C * H * W]`` or + ``[C, H, W]``). For each missed class-major index: its heat, the strongest same-class cell of its ``kernel x + kernel`` window that the DEVICE selected, and the relative gap ``(heat_missed - heat_neighbour) / heat_missed`` + (a few 0.1 % = a local-max near-tie: the neighbour won the device's 3x3 comparison; a bf16 heat holds 3 + significant digits). ``neighbour`` is None when no same-class neighbour was selected.""" + heat = np.asarray(reference_heat, np.float64).reshape(-1) + cells = int(height) * int(width) + sel = set(np.asarray(device_indices).astype(np.int64).reshape(-1).tolist()) + r = int(kernel) // 2 + out: List[Dict[str, Any]] = [] + for idx in missed: + c, p = divmod(int(idx), cells) + y, x = divmod(p, int(width)) + best = None + for dy in range(-r, r + 1): + for dx in range(-r, r + 1): + yy, xx = y + dy, x + dx + if (dy or dx) and 0 <= yy < height and 0 <= xx < width: + q = c * cells + yy * int(width) + xx + if q in sel and (best is None or heat[q] > heat[best]): + best = q + rec: Dict[str, Any] = {"index": int(idx), "class": c, "cell": [y, x], "heat": float(heat[idx]), + "neighbour": None if best is None else int(best)} + if best is not None: + rec["neighbour_heat"] = float(heat[best]) + rec["relative_gap"] = float((heat[idx] - heat[best]) / max(heat[idx], 1e-30)) + out.append(rec) + return out diff --git a/code/tt_diffusion_planner/ttaw/ops/upsample.py b/code/tt_diffusion_planner/ttaw/ops/upsample.py new file mode 100644 index 0000000000000000000000000000000000000000..9ebdc8b605c9a576df24b4a1df76ff72f757ac37 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/ops/upsample.py @@ -0,0 +1,203 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C18: up-sampling and interpolation helpers for NHWC feature maps (PLAN.md section 1.2 row C18). + +- :func:`upsample_nearest`: integer-factor nearest up-sampling through ``ttnn.upsample`` (floor indexing, exactly + ``F.interpolate(mode="nearest")`` and ONNX Resize nearest / asymmetric / floor: S:yolox section 4.3). ttnn's + nearest path takes ROW_MAJOR (interleaved or height / block sharded) or a TILE input whose W and C are multiples + of 32 (``upsample_common.cpp:25-65``) and always writes ROW_MAJOR, so the map goes TILE -> ROW_MAJOR -> + ``[N, H, W, C]`` -> upsample -> ``[1, 1, N*sH*H*sW*W, C]`` -> TILE (``output_layout``). Bit-exact (a data move). +- :func:`interp_matrix`: the 1-D interpolation matrix ``A [out, in]`` of ``F.interpolate`` (``"nearest"`` floor + rule, ``"linear"`` with ``align_corners`` True / False, the half-pixel rule clamped at 0 like torch), host numpy + float64. A 2-D resize is ``Y_c = A_h X_c A_w^T`` per channel (S:frnet:449, S:bevdet:488, S:meteor:497). +- :class:`Resize2d`: that separable resize on the device for any sizes (integer or not), e.g. the + ``align_corners=True`` bilinear of FRNet / BEVDet's FPN_LSS, which ``ttnn.upsample`` cannot do (its bilinear path + is half-pixel only, integer scales, sharded bf16 input). The W pass is a matmul over the transposed rows + (``[N*H, C, W] @ A_w^T``), the H pass ``A_h @ [H, W'*C]`` per image (``kron(I_N, A_h)`` for N > 1). Default + operands are float32 with HiFi4 + fp32 accumulation, because bf16 cannot hold weights such as 511/1023 (S:frnet:449) + and a device fp32 matmul is TF32-like (~1.5e-3 relative, probe P12); ``dtype="bfloat16"`` halves the traffic. + Static non-integer resizes may also use a precomputed ``grid_sample`` grid (probe P9: int16 corner indices exact, + bf16 weights; 1.7x faster than fp32 grids) through C23. +""" +from __future__ import annotations + +from typing import Any, Dict, Optional, Sequence, Tuple, Union + +import numpy as np + +from ..precision import Precision +from ..tensors import dtype_name, ttnn_dtype +from .conv import FeatureMap, _layout, _memory_config + +__all__ = ["upsample_nearest", "interp_matrix", "resize_matrices", "Resize2d"] + + +def _pair(value: Union[int, Sequence[int]]) -> Tuple[int, int]: + if isinstance(value, (int, np.integer)): + return int(value), int(value) + a, b = (int(x) for x in value) + return a, b + + +def upsample_nearest(fm: FeatureMap, scale: Union[int, Sequence[int]] = 2, *, output_layout: Any = "tile", + memory_config: Any = "dram") -> FeatureMap: + """Nearest up-sampling by integer factors (bit-exact). Returns ``[1, 1, N*sH*H*sW*W, C]`` in + ``output_layout`` (TILE by default, what a following conv / concat wants).""" + import ttnn + + sh, sw = _pair(scale) + if sh < 1 or sw < 1: + raise ValueError(f"scale must be positive integers, got {scale!r}") + t = fm.tensor + if t.layout != ttnn.ROW_MAJOR_LAYOUT: + t = ttnn.to_layout(t, ttnn.ROW_MAJOR_LAYOUT) + t = ttnn.reshape(t, (fm.batch, fm.height, fm.width, fm.channels)) + up = ttnn.upsample(t, (sh, sw), mode="nearest", memory_config=_memory_config(memory_config)) + out = ttnn.reshape(up, (1, 1, fm.batch * fm.height * sh * fm.width * sw, fm.channels)) + if _layout(output_layout) == ttnn.TILE_LAYOUT: + out = ttnn.to_layout(out, ttnn.TILE_LAYOUT) + return FeatureMap(out, fm.batch, fm.height * sh, fm.width * sw, fm.channels) + + +def interp_matrix(in_size: int, out_size: int, *, mode: str = "linear", align_corners: bool = False, + scale: Optional[float] = None) -> np.ndarray: + """``A`` (float64, ``[out_size, in_size]``) with ``y = A @ x`` equal to 1-D ``F.interpolate`` of ``x``. + + ``mode``: ``"nearest"`` (``src = min(floor(dst * s), in - 1)``) or ``"linear"`` (``align_corners=True``: + ``src = dst * (in - 1) / (out - 1)``; False: ``src = max((dst + 0.5) * s - 0.5, 0)``). ``s`` is ``in / out``, + or ``1 / scale`` when the caller passed ``scale_factor=scale`` to torch (its default + ``recompute_scale_factor=None`` keeps the given factor).""" + if in_size < 1 or out_size < 1: + raise ValueError(f"sizes must be positive, got {in_size} -> {out_size}") + s = (1.0 / float(scale)) if scale is not None else in_size / out_size + dst = np.arange(out_size, dtype=np.float64) + a = np.zeros((out_size, in_size), np.float64) + rows = np.arange(out_size) + if mode == "nearest": + if align_corners: + raise ValueError("align_corners applies to linear modes only") + # torch's nearest_idx: floorf(dst * (float)scale) in float32 (the exact 2x case equals dst >> 1) + prod = (dst.astype(np.float32) * np.float32(s)).astype(np.float32) + src = np.minimum(np.floor(prod).astype(np.int64), in_size - 1) + a[rows, src] = 1.0 + return a + if mode != "linear": + raise ValueError(f"mode {mode!r}: expected 'nearest' or 'linear'") + if align_corners: + src = dst * ((in_size - 1) / (out_size - 1)) if out_size > 1 else np.zeros(out_size) + else: + src = np.maximum((dst + 0.5) * s - 0.5, 0.0) + i0 = np.minimum(np.floor(src).astype(np.int64), in_size - 1) + i1 = np.minimum(i0 + 1, in_size - 1) + lam = src - i0 + np.add.at(a, (rows, i0), 1.0 - lam) + np.add.at(a, (rows, i1), lam) + return a + + +def resize_matrices(in_hw: Tuple[int, int], out_hw: Tuple[int, int], *, mode: str = "linear", + align_corners: bool = False, scale: Optional[Union[float, Sequence[float]]] = None + ) -> Tuple[np.ndarray, np.ndarray]: + """``(A_h [H', H], A_w [W', W])`` of a 2-D resize: ``Y_c = A_h @ X_c @ A_w.T``.""" + sh, sw = (None, None) if scale is None else (_pair_f(scale)) + return (interp_matrix(in_hw[0], out_hw[0], mode=mode, align_corners=align_corners, scale=sh), + interp_matrix(in_hw[1], out_hw[1], mode=mode, align_corners=align_corners, scale=sw)) + + +def _pair_f(value: Union[float, Sequence[float]]) -> Tuple[float, float]: + if isinstance(value, (int, float, np.floating, np.integer)): + return float(value), float(value) + a, b = (float(x) for x in value) + return a, b + + +class Resize2d: + """Separable 2-D resize of an NHWC feature map on the device (see the module docstring). + + Args: + in_hw, out_hw: input / output spatial size. ``channels``, ``batch``: map size (static: one instance per + call site). + mode, align_corners, scale: as :func:`interp_matrix`. + dtype: operand dtype of the two matmuls (``"float32"`` default; the input is typecast, the output cast + back to ``output_dtype``). + precision: compute config of the matmuls (``"accurate"``: HiFi4 + fp32 accumulation; never LoFi for fp32 + operands, probe P12). + Matrices are uploaded on the first call (warm-up), never inside a capture.""" + + def __init__(self, in_hw: Tuple[int, int], out_hw: Tuple[int, int], *, channels: int, batch: int = 1, + mode: str = "linear", align_corners: bool = False, scale: Optional[Any] = None, + dtype: Any = "float32", output_dtype: Any = "bfloat16", precision: Any = "accurate", + output_layout: Any = "tile", name: str = "resize"): + self.name = name + self.in_hw = (int(in_hw[0]), int(in_hw[1])) + self.out_hw = (int(out_hw[0]), int(out_hw[1])) + self.channels, self.batch = int(channels), int(batch) + self.mode, self.align_corners = mode, bool(align_corners) + self.a_h, self.a_w = resize_matrices(self.in_hw, self.out_hw, mode=mode, align_corners=align_corners, + scale=scale) + self.dtype = dtype_name(dtype) + self.output_dtype = dtype_name(output_dtype) + self.precision = Precision.parse(precision) + self.output_layout = output_layout + self._dev: Optional[Tuple[Any, Any]] = None + + def host_matrices(self) -> Tuple[np.ndarray, np.ndarray]: + """The device operands: ``A_w^T [W, W']`` and ``kron(I_N, A_h) [N*H', N*H]`` (float32).""" + aw_t = np.ascontiguousarray(self.a_w.T, np.float32) + ah = self.a_h if self.batch == 1 else np.kron(np.eye(self.batch), self.a_h) + return aw_t, np.ascontiguousarray(ah, np.float32) + + def reference(self, nchw: np.ndarray) -> np.ndarray: + """Host float64 result of the same matrices (tests).""" + x = np.asarray(nchw, np.float64) + return np.einsum("ph,nchw,qw->ncpq", self.a_h, x, self.a_w) + + def __call__(self, fm: FeatureMap) -> FeatureMap: + import torch + import ttnn + + n, (h, w), c = self.batch, self.in_hw, self.channels + (h2, w2) = self.out_hw + if (fm.batch, fm.height, fm.width, fm.channels) != (n, h, w, c): + got = (fm.batch, fm.height, fm.width, fm.channels) + raise ValueError(f"{self.name}: built for {(n, h, w, c)}, got {got}") + dt = ttnn_dtype(self.dtype) + cfg = self.precision.compute_kernel_config() + if self._dev is None: + aw_t, ah = self.host_matrices() + dev = fm.tensor.device() + self._dev = (ttnn.from_torch(torch.from_numpy(aw_t), dtype=dt, layout=ttnn.TILE_LAYOUT, device=dev), + ttnn.from_torch(torch.from_numpy(ah), dtype=dt, layout=ttnn.TILE_LAYOUT, device=dev)) + aw_t_dev, ah_dev = self._dev + x = fm.tensor + if x.layout != ttnn.ROW_MAJOR_LAYOUT: + x = ttnn.to_layout(x, ttnn.ROW_MAJOR_LAYOUT) + x = ttnn.reshape(x, (1, n * h, w, c)) + x = ttnn.to_layout(x, ttnn.TILE_LAYOUT) + if x.dtype != dt: + x = ttnn.typecast(x, dt) + xt = ttnn.transpose(x, -2, -1) # [1, N*H, C, W] + y = ttnn.matmul(xt, aw_t_dev, compute_kernel_config=cfg) # [1, N*H, C, W'] + y = ttnn.transpose(y, -2, -1) # [1, N*H, W', C] + y = ttnn.to_layout(y, ttnn.ROW_MAJOR_LAYOUT) + y = ttnn.reshape(y, (1, 1, n * h, w2 * c)) + y = ttnn.to_layout(y, ttnn.TILE_LAYOUT) + z = ttnn.matmul(ah_dev, y, compute_kernel_config=cfg) # [1, 1, N*H', W'*C] + if z.dtype != ttnn_dtype(self.output_dtype): + z = ttnn.typecast(z, ttnn_dtype(self.output_dtype)) # TILE typecast (every dtype pair) + z = ttnn.to_layout(z, ttnn.ROW_MAJOR_LAYOUT) + z = ttnn.reshape(z, (1, 1, n * h2 * w2, c)) + if _layout(self.output_layout) == ttnn.TILE_LAYOUT: + z = ttnn.to_layout(z, ttnn.TILE_LAYOUT) + return FeatureMap(z, n, h2, w2, c) + + def release(self) -> None: + import ttnn + + for t in self._dev or (): + if t.is_allocated(): + ttnn.deallocate(t) + self._dev = None + + def describe(self) -> Dict[str, Any]: + return {"name": self.name, "in_hw": list(self.in_hw), "out_hw": list(self.out_hw), "channels": self.channels, + "batch": self.batch, "mode": self.mode, "align_corners": self.align_corners, "dtype": self.dtype} diff --git a/code/tt_diffusion_planner/ttaw/outputs.py b/code/tt_diffusion_planner/ttaw/outputs.py new file mode 100644 index 0000000000000000000000000000000000000000..34e3c034edce2da9e38c88516edce8994ac222c8 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/outputs.py @@ -0,0 +1,283 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C08 outputs: the result classes of the 13 bundles and their ``POST /predict`` JSON (BUNDLE_CONVENTIONS.md 7.4). + +Every class has ``to_dicts()`` (one plain dict per detection / pose / class) and ``to_dict(output_format)`` (the +whole response body: ``model``, ``frame_id``, the payload, ``meta``, ``timing_ms``). The Python API returns these +objects and the server returns ``to_dict(...)`` of the same object, so both agree bit for bit. + +================== ================================= ================================================== +class families payload +================== ================================= ================================================== +``Detections3D`` CenterPoint, TransFusion, BEVFusion, ``detections`` [label, label_id, score, center, size, + PointPainting, StreamPETR, BEVDet, yaw, velocity], ``num_detections``; rows sorted by + BEVFormer, PTv3-det descending score at construction +``Detections2D`` YOLOX ``detections`` [label, label_id, score, box_xyxy] + + ``extras`` (e.g. ``{"semseg": Mask2D}``) +``Segmentation3D`` FRNet, PTv3-seg ``labels`` (npz uint8/uint16 [N]), ``scores``, + ``class_names``, ``class_counts`` +``Mask2D`` SceneSeg (YOLOX semseg extra) ``mask`` (png or npz), ``class_names``, ``class_counts`` +``Trajectory`` Diffusion Planner ``trajectory`` [[x, y, yaw, ...] x T], ``columns``, + ``turn_indicator``, ``predicted_agents`` +================== ================================= ================================================== +""" +from __future__ import annotations + +from dataclasses import dataclass, field +from typing import Any, Dict, List, Optional, Sequence + +import numpy as np + +from .io import encode_array, encode_png, to_jsonable + +__all__ = ["Detections3D", "Detections2D", "Segmentation3D", "Mask2D", "Trajectory", "label_name"] + + +def label_name(labels: Sequence[str], label_id: int) -> str: + """``labels[label_id]``, or the id as text when it is out of range.""" + return labels[label_id] if 0 <= label_id < len(labels) else str(label_id) + + +def _envelope(model: str, frame_id: str, meta: dict, timing_ms: dict) -> Dict[str, Any]: + return {"model": model, "frame_id": frame_id, "meta": to_jsonable(meta), + "timing_ms": {k: round(float(v), 3) for k, v in timing_ms.items()}} + + +def _sorted_by_score(scores: np.ndarray) -> np.ndarray: + return np.argsort(-scores, kind="stable") + + +@dataclass +class Detections3D: + """3D boxes in ``frame_id`` (Autoware ``base_link``: x forward, y left, z up; metres, radians). + + ``boxes`` float32 ``[N, 7]`` = x, y, z (box centre), length, width, height, yaw; ``scores`` ``[N]``; + ``label_ids`` int32 ``[N]`` (index into ``labels``); ``velocities`` ``[N, 2]`` (vx, vy) m/s or None. + Rows are sorted by descending score (stable) at construction.""" + + boxes: np.ndarray + scores: np.ndarray + label_ids: np.ndarray + velocities: Optional[np.ndarray] = None + labels: Sequence[str] = () + model: str = "" + frame_id: str = "base_link" + timing_ms: dict = field(default_factory=dict) + meta: dict = field(default_factory=dict) + + def __post_init__(self) -> None: + self.boxes = np.asarray(self.boxes, np.float32).reshape(-1, 7) + self.scores = np.asarray(self.scores, np.float32).reshape(-1) + self.label_ids = np.asarray(self.label_ids, np.int32).reshape(-1) + n = len(self.scores) + if len(self.boxes) != n or len(self.label_ids) != n: + raise ValueError(f"boxes {self.boxes.shape}, scores {self.scores.shape}, labels {self.label_ids.shape}") + if self.velocities is not None: + self.velocities = np.asarray(self.velocities, np.float32).reshape(n, 2) + order = _sorted_by_score(self.scores) + self.boxes, self.scores, self.label_ids = self.boxes[order], self.scores[order], self.label_ids[order] + if self.velocities is not None: + self.velocities = self.velocities[order] + + def __len__(self) -> int: + return int(self.scores.shape[0]) + + @property + def label_names(self) -> List[str]: + return [label_name(self.labels, i) for i in self.label_ids.tolist()] + + def to_dicts(self) -> List[Dict[str, Any]]: + out = [] + for i, name in enumerate(self.label_names): + x, y, z, length, width, height, yaw = (float(v) for v in self.boxes[i]) + d = {"label": name, "label_id": int(self.label_ids[i]), "score": round(float(self.scores[i]), 4), + "center": [round(x, 3), round(y, 3), round(z, 3)], + "size": [round(length, 3), round(width, 3), round(height, 3)], "yaw": round(yaw, 4)} + if self.velocities is not None: + d["velocity"] = [round(float(v), 3) for v in self.velocities[i]] + out.append(d) + return out + + def to_dict(self, output_format: str = "json") -> Dict[str, Any]: + body = _envelope(self.model, self.frame_id, self.meta, self.timing_ms) + body.update(num_detections=len(self), detections=self.to_dicts()) + if output_format == "npz": # lossless arrays for programmatic clients + body["arrays"] = {"boxes": encode_array(self.boxes, key="boxes"), + "scores": encode_array(self.scores, key="scores"), + "label_ids": encode_array(self.label_ids, key="label_ids")} + if self.velocities is not None: + body["arrays"]["velocities"] = encode_array(self.velocities, key="velocities") + return body + + +@dataclass +class Mask2D: + """A per-pixel class map (H, W) (uint8 for <= 256 classes) in the source image's pixel grid.""" + + mask: np.ndarray + class_names: Sequence[str] = () + model: str = "" + frame_id: str = "camera" + encoding: str = "png" + timing_ms: dict = field(default_factory=dict) + meta: dict = field(default_factory=dict) + + def __post_init__(self) -> None: + self.mask = np.asarray(self.mask) + if self.mask.ndim != 2: + raise ValueError(f"mask must be (H, W), got {self.mask.shape}") + if self.encoding not in ("png", "npz"): + raise ValueError("encoding must be 'png' or 'npz'") + + def class_counts(self) -> Dict[str, int]: + ids, counts = np.unique(self.mask, return_counts=True) + return {label_name(self.class_names, int(i)): int(c) for i, c in zip(ids, counts)} + + def to_dicts(self) -> List[Dict[str, Any]]: + ids, counts = np.unique(self.mask, return_counts=True) + return [{"label": label_name(self.class_names, int(i)), "label_id": int(i), "pixels": int(c)} + for i, c in zip(ids, counts)] + + def payload(self, output_format: str = "json") -> Dict[str, Any]: + """The encoded mask alone (also used when the mask is an extra of another output).""" + if self.encoding == "png" and output_format != "npz" and self.mask.dtype == np.uint8: + return encode_png(self.mask, key="mask") + return encode_array(self.mask, key="mask") + + def to_dict(self, output_format: str = "json") -> Dict[str, Any]: + body = _envelope(self.model, self.frame_id, self.meta, self.timing_ms) + body.update(mask=self.payload(output_format), class_names=list(self.class_names), + class_counts=self.class_counts()) + return body + + +@dataclass +class Detections2D: + """2-D boxes in original image pixels: ``boxes_xyxy`` float32 ``[N, 4]``, ``scores``, ``label_ids``. + ``extras`` holds companion outputs encoded into the body by name (``{"semseg": Mask2D(...)}``). + Rows are sorted by descending score (stable) at construction.""" + + boxes_xyxy: np.ndarray + scores: np.ndarray + label_ids: np.ndarray + labels: Sequence[str] = () + extras: Dict[str, Any] = field(default_factory=dict) + model: str = "" + frame_id: str = "camera" + timing_ms: dict = field(default_factory=dict) + meta: dict = field(default_factory=dict) + + def __post_init__(self) -> None: + self.boxes_xyxy = np.asarray(self.boxes_xyxy, np.float32).reshape(-1, 4) + self.scores = np.asarray(self.scores, np.float32).reshape(-1) + self.label_ids = np.asarray(self.label_ids, np.int32).reshape(-1) + if not (len(self.boxes_xyxy) == len(self.scores) == len(self.label_ids)): + raise ValueError("boxes_xyxy, scores and label_ids lengths differ") + order = _sorted_by_score(self.scores) + self.boxes_xyxy, self.scores, self.label_ids = self.boxes_xyxy[order], self.scores[order], self.label_ids[order] + + def __len__(self) -> int: + return int(self.scores.shape[0]) + + def to_dicts(self) -> List[Dict[str, Any]]: + return [{"label": label_name(self.labels, int(self.label_ids[i])), "label_id": int(self.label_ids[i]), + "score": round(float(self.scores[i]), 4), + "box_xyxy": [round(float(v), 2) for v in self.boxes_xyxy[i]]} for i in range(len(self))] + + def to_dict(self, output_format: str = "json") -> Dict[str, Any]: + body = _envelope(self.model, self.frame_id, self.meta, self.timing_ms) + body.update(num_detections=len(self), detections=self.to_dicts()) + for name, extra in self.extras.items(): + if hasattr(extra, "payload"): + body[name] = extra.payload(output_format) + elif isinstance(extra, np.ndarray): + body[name] = encode_array(extra, key=name) + else: + body[name] = to_jsonable(extra) + if output_format == "npz": + body["arrays"] = {"boxes_xyxy": encode_array(self.boxes_xyxy, key="boxes_xyxy"), + "scores": encode_array(self.scores, key="scores"), + "label_ids": encode_array(self.label_ids, key="label_ids")} + return body + + +@dataclass +class Segmentation3D: + """Per-point classes, in input point order after NaN removal (stated in SERVING.md). ``label_ids`` ``[N]``; + ``scores`` optional ``[N]`` (winning-class probability) or ``[N, C]``.""" + + label_ids: np.ndarray + class_names: Sequence[str] = () + scores: Optional[np.ndarray] = None + model: str = "" + frame_id: str = "base_link" + timing_ms: dict = field(default_factory=dict) + meta: dict = field(default_factory=dict) + + def __post_init__(self) -> None: + ids = np.asarray(self.label_ids).reshape(-1) + if ids.size and ids.min() < 0: + raise ValueError("label ids must be >= 0") + self.label_ids = ids.astype(np.uint8 if (not ids.size or ids.max() < 256) else np.uint16) + if self.scores is not None: + self.scores = np.asarray(self.scores, np.float32) + if self.scores.shape[0] != ids.shape[0]: + raise ValueError("scores and label_ids lengths differ") + + def __len__(self) -> int: + return int(self.label_ids.shape[0]) + + def class_counts(self) -> Dict[str, int]: + counts = np.bincount(self.label_ids.astype(np.int64), minlength=len(self.class_names)) + return {label_name(self.class_names, i): int(c) for i, c in enumerate(counts) if c or i < len(self.class_names)} + + def to_dicts(self) -> List[Dict[str, Any]]: + counts = np.bincount(self.label_ids.astype(np.int64), minlength=len(self.class_names)) + return [{"label": label_name(self.class_names, i), "label_id": i, "points": int(c)} + for i, c in enumerate(counts)] + + def to_dict(self, output_format: str = "json") -> Dict[str, Any]: + body = _envelope(self.model, self.frame_id, self.meta, self.timing_ms) + body.update(num_points=len(self), labels=encode_array(self.label_ids, key="labels"), + class_names=list(self.class_names), class_counts=self.class_counts()) + if self.scores is not None: + body["scores"] = encode_array(self.scores, key="scores") + return body + + +@dataclass +class Trajectory: + """A planned trajectory ``poses`` ``[T, D]`` whose columns are named by ``columns`` (default x, y, yaw) in + ``frame_id``; optional ``turn_indicator`` and ``predicted_agents`` ``[A, T, D']``.""" + + poses: np.ndarray + columns: Sequence[str] = ("x", "y", "yaw") + turn_indicator: Any = None + predicted_agents: Optional[np.ndarray] = None + model: str = "" + frame_id: str = "base_link" + timing_ms: dict = field(default_factory=dict) + meta: dict = field(default_factory=dict) + + def __post_init__(self) -> None: + self.poses = np.asarray(self.poses, np.float32) + if self.poses.ndim != 2 or self.poses.shape[1] != len(self.columns): + raise ValueError(f"poses {self.poses.shape} do not match columns {list(self.columns)}") + if self.predicted_agents is not None: + self.predicted_agents = np.asarray(self.predicted_agents, np.float32) + + def __len__(self) -> int: + return int(self.poses.shape[0]) + + def to_dicts(self) -> List[Dict[str, Any]]: + return [{c: round(float(v), 4) for c, v in zip(self.columns, row)} for row in self.poses] + + def to_dict(self, output_format: str = "json") -> Dict[str, Any]: + body = _envelope(self.model, self.frame_id, self.meta, self.timing_ms) + body.update(num_poses=len(self), columns=list(self.columns), + trajectory=[[round(float(v), 4) for v in row] for row in self.poses], + turn_indicator=to_jsonable(self.turn_indicator)) + if self.predicted_agents is not None: + body["predicted_agents"] = encode_array(self.predicted_agents, key="predicted_agents") + if output_format == "npz": + body["arrays"] = {"poses": encode_array(self.poses, key="poses")} + return body diff --git a/code/tt_diffusion_planner/ttaw/pointcloud.py b/code/tt_diffusion_planner/ttaw/pointcloud.py new file mode 100644 index 0000000000000000000000000000000000000000..9d94aa12a3a017517b2a49cd4c7c6183cae1bd45 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/pointcloud.py @@ -0,0 +1,282 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C11 point clouds: Autoware's multi-sweep densification state machine, input hygiene and the seeded permutation. + +**Densification** (``PointCloudDensification`` + ``VoxelGenerator::generateSweepPoints`` of +``autoware_lidar_centerpoint``; the same code shape is in lidar_transfusion, bevfusion and +image_projection_based_fusion: S:centerpoint:164-176; S:transfusion:99-103; S:pointpainting:216-221): + +- a cache holds the current cloud and ``num_past_frames`` previous ones, newest first (``push_front`` / + ``pop_back``, ``pointcloud_densification.cpp:79-95``); +- with a cache size > 1 every enqueue needs the TF ``T_sensor_from_world`` at the cloud's stamp; when it is missing + the frame is **skipped and not cached** (``:66-71``). The node keeps the world->current transform and each entry's + past->world transform as ``Eigen::Affine3f`` (float32, inverted in float32: ``:47-51,82-86``); +- at the end of each frame every cached cloud is re-transformed with ``world2current(now) @ past2world(entry)`` + (float32 product) by the float32 CUDA kernel, and gets ``time_lag = float(t_now - t_entry)`` in seconds + (``voxel_generator.cpp:56-93``). There is no maximum-age check: after dropped frames the past sweep can be old; +- points are written current sweep first; a sweep that would exceed ``cloud_capacity`` is dropped **with every older + sweep** (``voxel_generator.cpp:70-76``, a ``break``); +- cold start: the first frame has a single sweep (all ``time_lag = 0``); ``num_past_frames = 0`` needs no TF at all. + +:class:`SweepDensifier` is that state machine for one stream, :class:`StreamDensifiers` a per-``stream.id`` dict with +LRU eviction (PLAN.md D16: host sweep caches are per id), and :func:`densify_sweeps` the stateless variant for a +client that sends its past sweeps with ``T_current_from_sweep`` and ``time_lag_s`` itself. + +**Hygiene.** Autoware's range test is written as a rejection (``x < min || x >= max || ...``), so a NaN coordinate +passes it (S:centerpoint:179); the ports drop non-finite rows explicitly (a documented deviation; ``ttaw.io.PointCloud`` +already does it at decode time). Organized clouds mark missing returns as (0, 0, 0): :func:`nonzero_rows` drops them, +as the research references did (``centerpoint_ref.py``), before any transform. + +**Seeded permutation** (S:centerpoint:381-389; PLAN.md D1): Autoware reads its points through a fixed +``std::shuffle`` permutation at a time-seeded random offset, so which <= 32 points an over-full pillar keeps is random. +The ports keep the statistics but are deterministic: ``numpy.random.default_rng(seed).permutation(n)`` before the +first-K-points rule (:func:`seeded_permutation`). + +numpy only; no side effects on import. +""" +from __future__ import annotations + +import collections +from dataclasses import dataclass, field +from typing import Any, Dict, List, Optional, Sequence, Tuple + +import numpy as np + +from .geometry import compose_f32, invert_affine_f32, invert_rigid, transform_points_f32 + +__all__ = [ + "DEFAULT_CLOUD_CAPACITY", + "SweepDensifier", + "StreamDensifiers", + "DensifyInfo", + "densify_sweeps", + "finite_rows", + "nonzero_rows", + "range_mask", + "seeded_permutation", +] + +DEFAULT_CLOUD_CAPACITY = 2_000_000 # centerpoint_common.param.yaml:3 (cloud_capacity) + + +@dataclass +class DensifyInfo: + """What one densification produced: per used sweep its point count and time lag, plus the dropped sweeps.""" + + num_points: int = 0 + sweeps: List[Dict[str, Any]] = field(default_factory=list) # [{"points", "time_lag_s", "age_index"}] + dropped_sweeps: int = 0 # cached sweeps dropped by the capacity rule + + def to_dict(self) -> Dict[str, Any]: + return {"num_points": self.num_points, "sweeps": list(self.sweeps), "dropped_sweeps": self.dropped_sweeps} + + +def _sweep_rows(xyz: np.ndarray, time_lag: np.float32, extra: Optional[np.ndarray], point_feature_size: int, + affine: Optional[np.ndarray]) -> np.ndarray: + """Rows of one sweep in the network layout: (x, y, z, time_lag) or (x, y, z, intensity, time_lag).""" + out = np.empty((len(xyz), point_feature_size), np.float32) + out[:, :3] = xyz if affine is None else transform_points_f32(xyz, affine) + if point_feature_size == 4: + out[:, 3] = time_lag + elif point_feature_size == 5: + out[:, 3] = 0.0 if extra is None else extra + out[:, 4] = time_lag + else: + raise ValueError(f"point_feature_size must be 4 or 5, got {point_feature_size}") + return out + + +@dataclass +class _Cached: + xyz: np.ndarray # (N, 3) float32 in the frame of its own stamp + extra: Optional[np.ndarray] # (N,) float32 intensity (point_feature_size 5) or None + stamp_s: float + past2world: np.ndarray # (4, 4) float32 = Affine3f inverse of world2current at its stamp + + +class SweepDensifier: + """One stream's Autoware densification cache (module docstring). ``num_past_frames`` is + ``densification_params.num_past_frames`` (1 in Autoware's CenterPoint / TransFusion configs).""" + + def __init__(self, num_past_frames: int = 1, cloud_capacity: int = DEFAULT_CLOUD_CAPACITY): + if num_past_frames < 0: + raise ValueError("num_past_frames must be >= 0") + self.num_past_frames = int(num_past_frames) + self.cloud_capacity = int(cloud_capacity) + self._cache: List[_Cached] = [] + self._world2current = np.eye(4, dtype=np.float32) + self._stamp: Optional[float] = None + + @property + def cache_size(self) -> int: + """``pointcloud_cache_size()`` = past frames + the current one.""" + return self.num_past_frames + 1 + + @property + def needs_pose(self) -> bool: + """True when enqueue needs ``T_world_from_sensor`` (Autoware looks up a TF only for cache size > 1).""" + return self.cache_size > 1 + + def __len__(self) -> int: + return len(self._cache) + + @property + def stamps(self) -> List[float]: + """Stamps of the cached sweeps, newest first.""" + return [c.stamp_s for c in self._cache] + + def reset(self) -> None: + self._cache.clear() + self._world2current = np.eye(4, dtype=np.float32) + self._stamp = None + + def enqueue(self, xyz: np.ndarray, stamp_s: float, T_world_from_sensor: Optional[np.ndarray] = None, *, + intensity: Optional[np.ndarray] = None) -> bool: + """``PointCloudDensification::enqueuePointCloud``: returns False, leaving the cache unchanged, when a pose is + needed and missing (Autoware's TF failure: the frame is skipped). ``xyz`` (N, >=3) is in the sensor frame + of ``stamp_s`` (base_link for Autoware's concatenated cloud); ``T_world_from_sensor`` is that frame's pose + in the world (``map``) frame, float64.""" + pts = np.ascontiguousarray(np.asarray(xyz, dtype=np.float32)[:, :3]) + extra = None if intensity is None else np.asarray(intensity, dtype=np.float32).reshape(-1) + if extra is not None and len(extra) != len(pts): + raise ValueError("intensity length differs from the point count") + if self.needs_pose: + if T_world_from_sensor is None: + return False + # tf2 lookupTransform(frame <- world) in double, then .cast() (pointcloud_densification.cpp:47-51) + world2current = invert_rigid(np.asarray(T_world_from_sensor, dtype=np.float64)).astype(np.float32) + else: + world2current = np.eye(4, dtype=np.float32) + self._world2current = world2current + self._stamp = float(stamp_s) + self._cache.insert(0, _Cached(pts, extra, float(stamp_s), invert_affine_f32(world2current))) + if len(self._cache) > self.cache_size: + self._cache.pop() + return True + + def sweep_points(self, *, point_feature_size: int = 4) -> Tuple[np.ndarray, DensifyInfo]: + """``VoxelGenerator::generateSweepPoints``: every cached sweep in the current frame, newest first, as + (N, point_feature_size) float32 rows (x, y, z[, intensity], time_lag).""" + info = DensifyInfo() + if not self._cache: + return np.zeros((0, point_feature_size), np.float32), info + parts: List[np.ndarray] = [] + n = 0 + for k, c in enumerate(self._cache): + if n + len(c.xyz) > self.cloud_capacity: + info.dropped_sweeps = len(self._cache) - k + break + affine = compose_f32(self._world2current, c.past2world) + time_lag = np.float32(self._stamp - c.stamp_s) # double difference cast to float + parts.append(_sweep_rows(c.xyz, time_lag, c.extra, point_feature_size, affine)) + info.sweeps.append({"points": int(len(c.xyz)), "time_lag_s": float(time_lag), "age_index": k}) + n += len(c.xyz) + info.num_points = n + out = np.concatenate(parts, axis=0) if parts else np.zeros((0, point_feature_size), np.float32) + return out, info + + def describe(self) -> Dict[str, Any]: + return {"num_past_frames": self.num_past_frames, "cloud_capacity": self.cloud_capacity, + "cached": len(self._cache), "stamps": self.stamps} + + +class StreamDensifiers: + """Per-stream densification caches (PLAN.md D16: host sweep caches are a per-``stream.id`` dict). A new id beyond + ``max_streams`` evicts the least recently used stream; ``get(id, reset=True)`` restarts one (cold start).""" + + def __init__(self, num_past_frames: int = 1, cloud_capacity: int = DEFAULT_CLOUD_CAPACITY, + max_streams: int = 16): + if max_streams < 1: + raise ValueError("max_streams must be >= 1") + self.num_past_frames = int(num_past_frames) + self.cloud_capacity = int(cloud_capacity) + self.max_streams = int(max_streams) + self._streams: "collections.OrderedDict[str, SweepDensifier]" = collections.OrderedDict() + + def get(self, stream_id: str = "default", *, reset: bool = False) -> SweepDensifier: + sid = str(stream_id) + d = self._streams.get(sid) + if d is None: + d = SweepDensifier(self.num_past_frames, self.cloud_capacity) + self._streams[sid] = d + while len(self._streams) > self.max_streams: + self._streams.popitem(last=False) + elif reset: + d.reset() + self._streams.move_to_end(sid) + return d + + def forget(self, stream_id: str) -> None: + self._streams.pop(str(stream_id), None) + + @property + def streams(self) -> List[str]: + """Stream ids, most recently used first.""" + return list(reversed(self._streams)) + + def describe(self) -> Dict[str, Any]: + return {"max_streams": self.max_streams, "num_past_frames": self.num_past_frames, + "streams": {k: v.describe() for k, v in reversed(self._streams.items())}} + + +def densify_sweeps(current_xyz: np.ndarray, sweeps: Sequence[Tuple[np.ndarray, float, Optional[np.ndarray]]] = (), + *, point_feature_size: int = 4, cloud_capacity: int = DEFAULT_CLOUD_CAPACITY, + current_intensity: Optional[np.ndarray] = None, + sweep_intensity: Optional[Sequence[Optional[np.ndarray]]] = None) -> Tuple[np.ndarray, DensifyInfo]: + """Stateless densification for clients that hold the past sweeps: the current sweep (identity, ``time_lag`` + 0) then each ``(xyz, time_lag_s, T_current_from_sweep)`` in the given order (newest first, as Autoware's cache), + transformed by the float32 sweep kernel (``T`` None = already in the current frame) with + ``time_lag = float32(time_lag_s)``; the capacity rule drops a sweep and every later one.""" + info = DensifyInfo() + cur = np.asarray(current_xyz, dtype=np.float32)[:, :3] + rows = [(cur, np.float32(0.0), current_intensity, None)] + for i, (xyz, lag, T) in enumerate(sweeps): + extra = None if sweep_intensity is None else sweep_intensity[i] + rows.append((np.asarray(xyz, dtype=np.float32)[:, :3], np.float32(lag), extra, + None if T is None else np.asarray(T, dtype=np.float64))) + parts = [] + n = 0 + for k, (xyz, lag, extra, T) in enumerate(rows): + if n + len(xyz) > cloud_capacity: + info.dropped_sweeps = len(rows) - k + break + parts.append(_sweep_rows(xyz, lag, extra, point_feature_size, None if T is None else T.astype(np.float32))) + info.sweeps.append({"points": int(len(xyz)), "time_lag_s": float(lag), "age_index": k}) + n += len(xyz) + info.num_points = n + out = np.concatenate(parts, axis=0) if parts else np.zeros((0, point_feature_size), np.float32) + return out, info + + +# -------------------------------------------------------------------------------------------------- hygiene + +def finite_rows(points: np.ndarray, columns: Sequence[int] = (0, 1, 2)) -> np.ndarray: + """Mask of rows whose ``columns`` are all finite.""" + p = np.asarray(points) + return np.isfinite(p[:, list(columns)]).all(axis=1) + + +def nonzero_rows(points: np.ndarray) -> np.ndarray: + """Mask of rows whose x, y, z are not all exactly zero (organized clouds mark missing returns as (0, 0, 0); the + research references drop them with ``|x| + |y| + |z| > 0``, which also drops non-finite rows).""" + p = np.asarray(points, dtype=np.float32) + return np.abs(p[:, :3]).sum(axis=1) > 0 + + +def range_mask(points: np.ndarray, range_min: Sequence[float], range_max: Sequence[float]) -> np.ndarray: + """Autoware's half-open box test ``min <= v < max`` per axis on x, y, z (float32 comparisons). Unlike Autoware's + rejection form, a NaN coordinate fails it.""" + p = np.asarray(points, dtype=np.float32) + lo = np.asarray(range_min, dtype=np.float32) + hi = np.asarray(range_max, dtype=np.float32) + keep = np.ones(len(p), dtype=bool) + for a in range(3): + keep &= (p[:, a] >= lo[a]) & (p[:, a] < hi[a]) + return keep + + +def seeded_permutation(n: int, seed: Optional[int]) -> Optional[np.ndarray]: + """The deterministic stand-in for Autoware's random point shuffle: ``default_rng(seed).permutation(n)`` (PCG64, + numpy >= 1.17), or None when ``seed`` is None (input order).""" + if seed is None: + return None + return np.random.default_rng(int(seed)).permutation(int(n)) diff --git a/code/tt_diffusion_planner/ttaw/precision.py b/code/tt_diffusion_planner/ttaw/precision.py new file mode 100644 index 0000000000000000000000000000000000000000..ec51f4900813c43baf473b99287b37f18f8e6b0f --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/precision.py @@ -0,0 +1,191 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C05: explicit compute-kernel configs and per-module precision policies. + +Silent defaults this module exists to avoid (TT_PLATFORM.md section 0 item 7, PLAN.md section 0.2): + +- ``ttnn.matmul`` / ``ttnn.linear`` fall back to **LoFi** when a ``program_config`` or ``core_grid`` is given without + a ``compute_kernel_config``; +- ``ttnn.WormholeComputeKernelConfig()`` built without ``math_fidelity`` carries ``MathFidelity.Invalid``. + +So every op that takes a compute config gets one built here, with the fidelity spelled out. The default precision +is HiFi2 + fp32 accumulation, no approximations (the reference bundles' safe start; LoFi failed their gates almost +everywhere, RP section 2.9). A :class:`PrecisionPolicy` maps module names (globs, first match wins) to a +:class:`Precision`; ``_PRECISION`` overrides rules at build time (the A/B switch), e.g. +``CENTERPOINT_PRECISION="backbone.*=HiFi4+fp32;head.*=LoFi:w=bfp8"``. + +Other silent defaults to override by hand (not compute configs): ``ttnn.layer_norm`` epsilon 1e-12, SDPA +``is_causal=True``, ``ttnn.embedding`` PADDED returning the cached pad row, fused HARDSWISH skipped in conv2d. +""" +from __future__ import annotations + +import fnmatch +import os +from dataclasses import asdict, dataclass, replace +from typing import Any, Dict, List, Mapping, Optional, Sequence, Tuple, Union + +from .tensors import dtype_name + +__all__ = ["FIDELITIES", "Precision", "PRESETS", "compute_kernel_config", "PrecisionPolicy"] + +FIDELITIES = ("LoFi", "HiFi2", "HiFi3", "HiFi4") +_FIDELITY_KEY = {f.lower(): f for f in FIDELITIES} + + +def _fidelity(name: str) -> str: + key = _FIDELITY_KEY.get(str(name).strip().lower()) + if key is None: + raise ValueError(f"math fidelity {name!r}: expected one of {FIDELITIES}") + return key + + +def compute_kernel_config(fidelity: str = "HiFi2", *, fp32_acc: bool = True, approx: bool = False, + packer_l1_acc: bool = False, dst_full_sync: bool = False): + """A ``ttnn.WormholeComputeKernelConfig`` (the same class as ``BlackholeComputeKernelConfig``) with every + field explicit. ``fp32_acc`` = ``fp32_dest_acc_en`` (halves DST capacity: 4 tiles in half-sync).""" + import ttnn + + return ttnn.WormholeComputeKernelConfig(math_fidelity=getattr(ttnn.MathFidelity, _fidelity(fidelity)), + math_approx_mode=bool(approx), fp32_dest_acc_en=bool(fp32_acc), + packer_l1_acc=bool(packer_l1_acc), dst_full_sync_en=bool(dst_full_sync)) + + +@dataclass(frozen=True) +class Precision: + """Fidelity / accumulation / dtype choice of one module (an op or a group of ops).""" + + fidelity: str = "HiFi2" + fp32_acc: bool = True + approx: bool = False + packer_l1_acc: bool = False + dst_full_sync: bool = False + weights: str = "bfloat16" + activations: str = "bfloat16" + + def __post_init__(self) -> None: + object.__setattr__(self, "fidelity", _fidelity(self.fidelity)) + object.__setattr__(self, "weights", dtype_name(self.weights)) + object.__setattr__(self, "activations", dtype_name(self.activations)) + + @classmethod + def parse(cls, spec: Union[str, "Precision"]) -> "Precision": + """``"HiFi4+fp32"``, ``"LoFi"``, ``"HiFi2+fp32+approx+l1acc:w=bfp8:a=bf16"`` or a preset name + (``accurate`` / ``balanced`` / ``fast``). Flags: ``fp32`` (fp32 accumulation; absent = bf16 DST), + ``approx``, ``l1acc``, ``fullsync``; ``w=`` / ``a=`` set weight / activation dtypes.""" + if isinstance(spec, Precision): + return spec + text = spec.strip() + if text.lower() in PRESETS: + return PRESETS[text.lower()] + head, *opts = text.split(":") + fid, *flags = [p.strip() for p in head.split("+")] + kw: Dict[str, Any] = {"fidelity": fid, "fp32_acc": False} + for flag in (f.lower() for f in flags): + if flag == "fp32": + kw["fp32_acc"] = True + elif flag == "approx": + kw["approx"] = True + elif flag == "l1acc": + kw["packer_l1_acc"] = True + elif flag == "fullsync": + kw["dst_full_sync"] = True + else: + raise ValueError(f"precision {spec!r}: unknown flag {flag!r}") + for opt in opts: + k, _, v = opt.partition("=") + k = k.strip().lower() + if k in ("w", "weights"): + kw["weights"] = v.strip() + elif k in ("a", "act", "activations"): + kw["activations"] = v.strip() + else: + raise ValueError(f"precision {spec!r}: unknown option {k!r}") + return cls(**kw) + + @property + def label(self) -> str: + """Round-trips through :meth:`parse`.""" + flags = "".join(f"+{f}" for f, on in (("fp32", self.fp32_acc), ("approx", self.approx), + ("l1acc", self.packer_l1_acc), ("fullsync", self.dst_full_sync)) if on) + return f"{self.fidelity}{flags}:w={self.weights}:a={self.activations}" + + def with_(self, **changes: Any) -> "Precision": + return replace(self, **changes) + + def compute_kernel_config(self): + return compute_kernel_config(self.fidelity, fp32_acc=self.fp32_acc, approx=self.approx, + packer_l1_acc=self.packer_l1_acc, dst_full_sync=self.dst_full_sync) + + def weights_dtype(self): + from .tensors import ttnn_dtype + + return ttnn_dtype(self.weights) + + def activations_dtype(self): + from .tensors import ttnn_dtype + + return ttnn_dtype(self.activations) + + def to_dict(self) -> Dict[str, Any]: + return asdict(self) + + +PRESETS: Dict[str, Precision] = { + "accurate": Precision("HiFi4", fp32_acc=True), # norms, softmax logits, box regression, grid sampling + "balanced": Precision("HiFi2", fp32_acc=True), # default for big matmuls / convs + "fast": Precision("LoFi", fp32_acc=False, weights="bfloat8_b"), # only after the gates pass with it +} + +RuleSpec = Union[Mapping[str, Union[str, Precision]], Sequence[Tuple[str, Union[str, Precision]]]] + + +class PrecisionPolicy: + """Ordered ``pattern -> Precision`` rules (``fnmatch`` globs on dotted module names, first match wins) plus a + default. Typical use, once at model build:: + + POLICY = PrecisionPolicy({"backbone.*": "balanced", "head.reg*": "accurate"}, default="balanced") + policy = POLICY.with_env("CENTERPOINT") # _PRECISION overrides, read once + cfg = policy.compute_kernel_config("backbone.block3.conv2") + w_dtype = policy.resolve("backbone.block3.conv2").weights_dtype() + + ``resolve`` records each module it answered for, so ``describe()`` shows the policy that reached the ops.""" + + def __init__(self, rules: Optional[RuleSpec] = None, default: Union[str, Precision] = "balanced"): + items = rules.items() if isinstance(rules, Mapping) else (rules or ()) + self.rules: List[Tuple[str, Precision]] = [(str(p), Precision.parse(v)) for p, v in items] + self.default = Precision.parse(default) + self.used: Dict[str, str] = {} + self._configs: Dict[Precision, Any] = {} + + def resolve(self, module: str) -> Precision: + for pattern, prec in self.rules: + if fnmatch.fnmatchcase(module, pattern): + self.used[module] = prec.label + return prec + self.used[module] = self.default.label + return self.default + + def compute_kernel_config(self, module: str): + prec = self.resolve(module) + if prec not in self._configs: + self._configs[prec] = prec.compute_kernel_config() + return self._configs[prec] + + def override(self, spec: str) -> "PrecisionPolicy": + """A new policy with ``"pattern=precision;pattern=precision"`` rules placed before the existing ones + (``*=...`` effectively replaces the default for unmatched modules).""" + extra = [] + for part in filter(None, (p.strip() for p in spec.split(";"))): + pattern, sep, prec = part.partition("=") + if not sep: + raise ValueError(f"precision override {part!r}: expected pattern=precision") + extra.append((pattern.strip(), Precision.parse(prec))) + return PrecisionPolicy(extra + self.rules, self.default) + + def with_env(self, prefix: str, env: Optional[Mapping[str, str]] = None) -> "PrecisionPolicy": + """Apply ``_PRECISION`` if set (else return ``self``). Call once at build.""" + spec = (os.environ if env is None else env).get(f"{prefix}_PRECISION", "").strip() + return self.override(spec) if spec else self + + def describe(self) -> Dict[str, Any]: + return {"default": self.default.label, "rules": [[p, prec.label] for p, prec in self.rules], + "used": dict(sorted(self.used.items()))} diff --git a/code/tt_diffusion_planner/ttaw/profiling.py b/code/tt_diffusion_planner/ttaw/profiling.py new file mode 100644 index 0000000000000000000000000000000000000000..66b2a2fef780aeccf1a151e36ae1ed4789e81c68 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/profiling.py @@ -0,0 +1,440 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C07: measurement helpers -- the numbers OPT_BASELINE.md / OPT_REPORT.md / the model card quote. + +- :class:`StageBench` collects wall times per stage (``host_in``, ``h2d``, ``trace``, ``b2b``, ``d2h``, ``post``, + ``e2e``) and reports p50 / p99 / mean / min / max in ms (RP section 3, rf-detr ``bench_breakdown``). +- :func:`bench_trace_runner` measures those stages on a :class:`.trace.TraceRunner` variant. +- :func:`signpost` / :func:`signposted` mark Tracy ranges (no-ops when the ``tracy`` module is absent), so + ``tt-perf-report --start-signpost trace --end-signpost trace_end`` and :func:`summarize_ops` can cut the CSV. +- :func:`summarize_ops` reads a device-profiler ``ops_perf_results*.csv`` (Tracy ``-r``) or the C++ + ``cpp_device_perf_report.csv`` and returns op count, kernel sum, op-to-op gaps, span, per-op-code totals and the + math-fidelity histogram. CLI: ``python -m .ttaw.profiling [--start trace] [--json out.json]``. +- :class:`AiclkSampler` samples the chip clock (sysfs ``tt_aiclk``) and hwmon power / temperature in a background + thread during a bench: AICLK sags from 1350 MHz under load (RP section 3), which explains span-vs-wall gaps. + +Profiling runs go through ``bin/devrun`` like any device job, with an absolute output directory under +``generated/profiler/_`` (PLAN.md section 5.1). +""" +from __future__ import annotations + +import argparse +import contextlib +import csv +import glob +import json +import logging +import os +import statistics +import sys +import threading +import time +from collections import defaultdict +from dataclasses import dataclass, field +from pathlib import Path +from typing import Any, Callable, Dict, Iterator, List, Mapping, Optional, Sequence, Union + +import numpy as np + +__all__ = [ + "STAGES", + "StageBench", + "time_b2b", + "bench_trace_runner", + "signpost", + "signposted", + "read_device_profiler", + "find_ops_csv", + "OpsSummary", + "summarize_ops", + "AiclkSampler", +] + +log = logging.getLogger(__name__) +STAGES = ("host_in", "h2d", "trace", "b2b", "d2h", "post", "e2e") +DEFAULT_AICLK_MHZ = 1350.0 + + +# ------------------------------------------------------------------------------------------- stage bench + +def _stats(samples: Sequence[float]) -> Dict[str, float]: + a = np.asarray(samples, np.float64) + return {"n": int(a.size), "p50": float(np.percentile(a, 50)), "p99": float(np.percentile(a, 99)), + "mean": float(a.mean()), "min": float(a.min()), "max": float(a.max())} + + +class StageBench: + """Per-stage wall-time samples in milliseconds. + + :: + + bench = StageBench("centerpoint eth/1cq") + for _ in range(100): + with bench.stage("e2e"): + with bench.stage("host_in"): host = prepare(points) + ... + print(bench.table()); bench.save_json("logs/centerpoint/bench_baseline.json", config=device_info) + """ + + def __init__(self, name: str = ""): + self.name = name + self.samples: Dict[str, List[float]] = {} + + @contextlib.contextmanager + def stage(self, name: str) -> Iterator[None]: + t0 = time.perf_counter() + try: + yield + finally: + self.add(name, (time.perf_counter() - t0) * 1e3) + + def add(self, name: str, ms: float) -> None: + self.samples.setdefault(name, []).append(float(ms)) + + def summary(self) -> Dict[str, Dict[str, float]]: + """``{stage: {n, p50, p99, mean, min, max}}`` in ms; known stages first, in pipeline order.""" + order = [s for s in STAGES if s in self.samples] + [s for s in self.samples if s not in STAGES] + return {s: _stats(self.samples[s]) for s in order if self.samples[s]} + + def table(self) -> str: + rows = [f"| stage | n | p50 ms | p99 ms | mean ms | min ms |", "|---|---:|---:|---:|---:|---:|"] + for s, st in self.summary().items(): + rows.append(f"| {s} | {st['n']} | {st['p50']:.3f} | {st['p99']:.3f} | {st['mean']:.3f} | {st['min']:.3f} |") + return (f"**{self.name}**\n\n" if self.name else "") + "\n".join(rows) + + def to_dict(self, **extra: Any) -> Dict[str, Any]: + return {"name": self.name, "stages_ms": self.summary(), **extra} + + def save_json(self, path: Union[str, os.PathLike], **extra: Any) -> Path: + path = Path(path) + path.parent.mkdir(parents=True, exist_ok=True) + path.write_text(json.dumps(self.to_dict(**extra), indent=1, default=str) + "\n") + return path + + +def time_b2b(enqueue: Callable[[], Any], sync: Callable[[], Any], n: int = 100, warmup: int = 5) -> float: + """Back-to-back device time per iteration (ms): ``warmup`` + ``n`` non-blocking ``enqueue()`` calls, one + ``sync()`` each round. With a traced ``enqueue`` this is the device time per forward.""" + for _ in range(warmup): + enqueue() + sync() + t0 = time.perf_counter() + for _ in range(n): + enqueue() + sync() + return (time.perf_counter() - t0) * 1e3 / n + + +def bench_trace_runner(runner, variant: Optional[str], inputs: Mapping[str, Any], *, + params: Optional[Mapping[str, Any]] = None, iters: int = 100, warmup: int = 10, + post: Optional[Callable[[Any], Any]] = None, b2b_iters: Optional[int] = None, + name: str = "") -> StageBench: + """Stage breakdown of one :class:`.trace.TraceRunner` variant (medians are the headline numbers): + + ``host_in`` host tensor conversion, ``h2d`` upload + sync, ``trace`` one replay + sync, ``d2h`` read, + ``post`` the optional ``post(outputs)`` callback, ``e2e`` ``runner(variant, inputs)`` + ``post``, and ``b2b`` + back-to-back replays (device time per forward).""" + import ttnn + + from .tensors import to_host_tensor + + dev = runner.device + bench = StageBench(name or f"{runner.name}/{variant}") + slots = {k: runner._slot(k, "input") for k in inputs} + for _ in range(warmup): + out = runner(variant, inputs, params) + if post is not None: + post(out) + for _ in range(iters): + t0 = time.perf_counter() + host = {k: to_host_tensor(v, slots[k].dtype, slots[k].layout, shape=slots[k].shape) for k, v in inputs.items()} + t1 = time.perf_counter() + runner.upload(host, params) + ttnn.synchronize_device(dev) + t2 = time.perf_counter() + runner.replay(variant) + ttnn.synchronize_device(dev) + t3 = time.perf_counter() + out = runner.read(variant) + t4 = time.perf_counter() + if post is not None: + post(out) + t5 = time.perf_counter() + for stage, a, b in (("host_in", t0, t1), ("h2d", t1, t2), ("trace", t2, t3), ("d2h", t3, t4)): + bench.add(stage, (b - a) * 1e3) + if post is not None: + bench.add("post", (t5 - t4) * 1e3) + for _ in range(iters): + with bench.stage("e2e"): + out = runner(variant, inputs, params) + if post is not None: + post(out) + bench.add("b2b", time_b2b(lambda: runner.replay(variant), lambda: ttnn.synchronize_device(dev), + n=b2b_iters or iters)) + return bench + + +# --------------------------------------------------------------------------------------------- signposts + +def signpost(name: str, message: Optional[str] = None) -> bool: + """Emit a Tracy signpost (a row in the ops CSV); returns False (no-op) when ``tracy`` is not importable.""" + try: + from tracy import signpost as _signpost + except ImportError: + return False + _signpost(header=name, message=message) + return True + + +@contextlib.contextmanager +def signposted(name: str) -> Iterator[None]: + """``signpost(name)`` ... ``signpost(name + "_end")`` around the block.""" + signpost(name) + try: + yield + finally: + signpost(f"{name}_end") + + +def read_device_profiler(device) -> bool: + """Flush the device profiler buffer (``ttnn.ReadDeviceProfiler``); call it per trace segment, the buffer holds + about 1000 ops. Returns False when the API is absent.""" + import ttnn + + fn = getattr(ttnn, "ReadDeviceProfiler", None) + if fn is None: + return False + fn(device) + return True + + +# ------------------------------------------------------------------------------------- ops CSV summary + +def find_ops_csv(path: Union[str, os.PathLike]) -> Path: + """A CSV file as-is, or the newest ``ops_perf_results*.csv`` (else ``cpp_device_perf_report.csv``) below a dir.""" + p = Path(path) + if p.is_file(): + return p + for pattern in ("ops_perf_results*.csv", "cpp_device_perf_report.csv"): + found = sorted(glob.glob(str(p / "**" / pattern), recursive=True), key=os.path.getmtime) + if found: + return Path(found[-1]) + raise FileNotFoundError(f"no ops_perf_results*.csv or cpp_device_perf_report.csv under {p}") + + +def _num(row: Mapping[str, str], key: str) -> float: + value = (row.get(key) or "").strip() + try: + return float(value) if value else 0.0 + except ValueError: + return 0.0 + + +def _chip_freq_mhz(csv_path: Path) -> Optional[float]: + """CHIP_FREQ[MHz] from a ``profile_log_device.csv`` header near the ops CSV, if any.""" + for cand in [csv_path.parent / "profile_log_device.csv", csv_path.parent.parent / "profile_log_device.csv", + csv_path.parent / ".logs" / "profile_log_device.csv"]: + try: + head = cand.read_text(errors="ignore").splitlines()[0] + except (OSError, IndexError): + continue + for part in head.split(","): + if "CHIP_FREQ" in part and ":" in part: + try: + return float(part.split(":")[1].strip()) + except ValueError: + pass + return None + + +@dataclass +class OpsSummary: + """Summary of one profiled section (times in microseconds).""" + + source: str + section: str + ops: int + kernel_sum_us: float + fw_sum_us: float + op2op_sum_us: float + span_us: Optional[float] + freq_mhz: float + by_op: List[Dict[str, Any]] = field(default_factory=list) + fidelity: Dict[str, int] = field(default_factory=dict) + top_gaps: List[Dict[str, Any]] = field(default_factory=list) + + def to_dict(self) -> Dict[str, Any]: + return {k: getattr(self, k) for k in self.__dataclass_fields__} + + def table(self, top: int = 20) -> str: + span = "n/a" if self.span_us is None else f"{self.span_us:.1f}" + lines = [f"{self.source} [{self.section}]", + f"ops {self.ops} kernel_sum {self.kernel_sum_us:.1f} us op2op_sum {self.op2op_sum_us:.1f} us " + f"span {span} us (@{self.freq_mhz:.0f} MHz) fidelity {self.fidelity}", + f"{'op code':55s} {'count':>6s} {'kernel us':>11s} {'share':>7s}"] + for row in self.by_op[:top]: + lines.append(f"{row['op'][:55]:55s} {row['count']:6d} {row['kernel_us']:11.1f} {row['share']:6.1%}") + return "\n".join(lines) + + +def summarize_ops(source: Union[str, os.PathLike, Sequence[Mapping[str, str]]], *, start: Optional[str] = "trace", + end: Optional[str] = None, last_replay_session: Optional[bool] = None, + freq_mhz: Optional[float] = None, top_gaps: int = 10) -> OpsSummary: + """Summarize device ops between signposts ``start`` and ``end`` (default: the next signpost). + + ``source``: an ``ops_perf_results*.csv`` / ``cpp_device_perf_report.csv`` path, a directory holding one, or + already-parsed rows. ``start=None`` takes every op. ``last_replay_session`` (default: on when the rows carry + ``METAL TRACE REPLAY SESSION ID`` values) keeps only the last trace replay session, the way ``prof_cpp.py`` + did. Span uses the FW start/end cycles at ``freq_mhz`` (default: the CHIP_FREQ of ``profile_log_device.csv``, + else 1350 MHz; AICLK sags under load, so treat span as approximate).""" + if isinstance(source, (str, os.PathLike)): + path = find_ops_csv(source) + with open(path, newline="") as f: + rows = list(csv.DictReader(f)) + label = str(path) + freq = freq_mhz or _chip_freq_mhz(path) or DEFAULT_AICLK_MHZ + else: + rows, label, freq = list(source), "", freq_mhz or DEFAULT_AICLK_MHZ + name_key = "OP CODE" if rows and "OP CODE" in rows[0] else "OP NAME" + section = "all" + if start is not None and any((r.get("OP TYPE") or "") == "signpost" for r in rows): + selected, on = [], False + for r in rows: + if (r.get("OP TYPE") or "") == "signpost": + code = r.get(name_key, "") + if on and (end is None or code == end): + break + on = on or code == start + continue + if on: + selected.append(r) + rows, section = selected, f"{start}..{end or 'next signpost'}" + else: + rows = [r for r in rows if (r.get("OP TYPE") or "") != "signpost"] + sessions = [r.get("METAL TRACE REPLAY SESSION ID", "") for r in rows] + if last_replay_session is None: + last_replay_session = any(s.strip() for s in sessions) + if last_replay_session and rows: + traced = [r for r in rows if (r.get("METAL TRACE ID") or "").strip()] + if traced: + last = traced[-1].get("METAL TRACE REPLAY SESSION ID") + rows = [r for r in traced if r.get("METAL TRACE REPLAY SESSION ID") == last] + section += f" (replay session {last})" + kernel = [_num(r, "DEVICE KERNEL DURATION [ns]") / 1e3 for r in rows] + fw = [_num(r, "DEVICE FW DURATION [ns]") / 1e3 for r in rows] + gaps = [_num(r, "OP TO OP LATENCY [ns]") / 1e3 for r in rows] + starts = [_num(r, "DEVICE FW START CYCLE") for r in rows] + ends = [_num(r, "DEVICE FW END CYCLE") for r in rows] + span = (max(ends) - min(s for s in starts if s > 0)) / freq if rows and any(starts) and any(ends) else None + agg: Dict[str, List[float]] = defaultdict(lambda: [0, 0.0]) + fidelity: Dict[str, int] = defaultdict(int) + for r, k in zip(rows, kernel): + agg[r.get(name_key, "?")][0] += 1 + agg[r.get(name_key, "?")][1] += k + fid = (r.get("MATH FIDELITY") or "").strip() + if fid: + fidelity[fid] += 1 + total = sum(kernel) or 1.0 + by_op = [{"op": op, "count": int(c), "kernel_us": t, "share": t / total} + for op, (c, t) in sorted(agg.items(), key=lambda kv: -kv[1][1])] + order = sorted(range(1, len(rows)), key=lambda i: -gaps[i])[:top_gaps] + worst = [{"index": i, "op": rows[i].get(name_key, "?"), "gap_us": gaps[i], "after": rows[i - 1].get(name_key, "?")} + for i in order] + return OpsSummary(label, section, len(rows), sum(kernel), sum(fw), sum(gaps[1:]), span, freq, by_op, + dict(fidelity), worst) + + +# ------------------------------------------------------------------------------------------------ AICLK + +class AiclkSampler: + """Background sampling of ``tt_aiclk`` (MHz) and hwmon power (W) / temperature (C) of one chip. + + :: + + with AiclkSampler(chip=0, interval_s=0.05) as clk: + run_bench() + print(clk.summary()) # {"aiclk_mhz": {"median": 1302, "min": 1206, ...}, "power_w": {...}, ...} + + Reads sysfs only (read-only, works while another process holds the device); ``available`` is False on hosts + without the Tenstorrent KMD.""" + + def __init__(self, chip: int = 0, interval_s: float = 0.05, *, root: str = "/sys/class/tenstorrent"): + self.interval_s = float(interval_s) + base = Path(root) / f"tenstorrent!{chip}" + self._aiclk = base / "tt_aiclk" + hwmon = sorted(glob.glob(str(base / "device" / "hwmon" / "hwmon*"))) + self._hwmon = Path(hwmon[0]) if hwmon else None + self.available = self._aiclk.is_file() + self.samples: Dict[str, List[float]] = {"aiclk_mhz": [], "power_w": [], "temp_c": []} + self._stop = threading.Event() + self._thread: Optional[threading.Thread] = None + + @staticmethod + def _read(path: Optional[Path], scale: float) -> Optional[float]: + if path is None: + return None + try: + return float(path.read_text().split()[0]) * scale + except (OSError, ValueError, IndexError): + return None + + def sample_once(self) -> Dict[str, Optional[float]]: + hw = self._hwmon + values = {"aiclk_mhz": self._read(self._aiclk, 1.0), + "power_w": self._read(hw / "power1_input" if hw else None, 1e-6), + "temp_c": self._read(hw / "temp1_input" if hw else None, 1e-3)} + for k, v in values.items(): + if v is not None: + self.samples[k].append(v) + return values + + def _loop(self) -> None: + while not self._stop.is_set(): + self.sample_once() + self._stop.wait(self.interval_s) + + def start(self) -> "AiclkSampler": + if self.available and self._thread is None: + self._stop.clear() + self._thread = threading.Thread(target=self._loop, name="aiclk-sampler", daemon=True) + self._thread.start() + return self + + def stop(self) -> None: + if self._thread is not None: + self._stop.set() + self._thread.join(timeout=5) + self._thread = None + + def __enter__(self) -> "AiclkSampler": + return self.start() + + def __exit__(self, *exc) -> None: + self.stop() + + def summary(self) -> Dict[str, Any]: + out: Dict[str, Any] = {"available": self.available} + for k, v in self.samples.items(): + if v: + out[k] = {"median": statistics.median(v), "min": min(v), "max": max(v), "n": len(v)} + return out + + +def main(argv: Optional[Sequence[str]] = None) -> int: + """``python -m .ttaw.profiling [--start trace] [--end trace_end] [--json out] [--top 25]``.""" + ap = argparse.ArgumentParser(description="Summarize a device-profiler ops CSV") + ap.add_argument("source") + ap.add_argument("--start", default="trace", help="signpost that opens the section ('' = all ops)") + ap.add_argument("--end", default=None) + ap.add_argument("--freq-mhz", type=float, default=None) + ap.add_argument("--top", type=int, default=25) + ap.add_argument("--json", default=None) + a = ap.parse_args(argv) + summary = summarize_ops(a.source, start=a.start or None, end=a.end, freq_mhz=a.freq_mhz) + print(summary.table(a.top)) + if a.json: + Path(a.json).write_text(json.dumps(summary.to_dict(), indent=1) + "\n") + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/code/tt_diffusion_planner/ttaw/server/__init__.py b/code/tt_diffusion_planner/ttaw/server/__init__.py new file mode 100644 index 0000000000000000000000000000000000000000..387839e47a6170bb57a15d7c4517389dfc34b02f --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/server/__init__.py @@ -0,0 +1,15 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C08 HTTP serving shared by every bundle. + +- ``server.app``: :func:`create_app` builds the FastAPI (ASGI) app of a bundle from a :class:`ServerSpec` + (routes ``/``, ``/health``, ``/v1/health``, ``/info``, ``/v1/models``, ``/predict``; lifespan = weights -> device -> + graph -> trace capture; errors 400 / 422 / 503 / 500). Imports fastapi / pydantic. +- ``server.client``: standard-library-only request builder, HTTP helpers and smoke checks (``/info`` must report + the pinned dispatch and grid). +- ``server.smoke``: standard-library-only agreement of a served body with a stored CPU-reference body (every output + family, numpy-free npz / png decoding) and the serve-profile pins of a staged tt-model manifest. + +Neither imports anything from ttaw, so a bundle's ``smoke_test.py`` can load them by path and run under any +Python 3.9+ outside the container. This ``__init__`` imports nothing, so ``import .ttaw.server.client`` never +pulls in fastapi. +""" diff --git a/code/tt_diffusion_planner/ttaw/server/app.py b/code/tt_diffusion_planner/ttaw/server/app.py new file mode 100644 index 0000000000000000000000000000000000000000..e8dc225a5ec6631cff0cc02651aae8ae23ad2665 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/server/app.py @@ -0,0 +1,401 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C08 server: the FastAPI app factory implementing the HTTP contract of all 13 bundles (BUNDLE_CONVENTIONS.md 7). + +A bundle's ``server/app.py`` is a few lines:: + + from ..api import CenterPoint + from ..ttaw.server.app import ServerSpec, create_app, parse_mesh_shape # noqa: F401 (verify: line) + + SPEC = ServerSpec(model_name=CenterPoint.MODEL_NAME, env_prefix="CENTERPOINT", model_cls=CenterPoint, + task="LiDAR 3D object detection", default_weights=CenterPoint.DEFAULT_REPO, + io="a LiDAR point cloud in, 3D boxes out", + autoware={"package": "autoware_lidar_centerpoint", + "path": "perception/autoware_lidar_centerpoint", + "autoware_universe": "9ceaccf026c31ffc5319bc9eeb4bd7bede0af3fd"}, + source={"repo": "https://huggingface.co/changh95/centerpoint-p150", "license": "Apache-2.0"}, + calib_dir=Path(__file__).resolve().parents[1] / "calib") + app = create_app(SPEC) + +Served by tt-model-manager as ``kind: tt-dit-server`` (``python -m uvicorn --lifespan on .server.app:app``). +Everything device-touching happens in the lifespan: weights -> device -> graph -> trace capture of every warm-up +variant; uvicorn logs ``Application startup complete`` only after that, so READY means warm. Shutdown (SIGTERM, +120 s budget) closes the model under the lock. Importing the module has no side effects beyond importing +fastapi / pydantic (no device, no download, no environment read). + +Tests replace the model with a stub: ``app.state.ttaw.model_factory = StubModel`` before ``TestClient(app)``. +``app.state.ttaw.predict(request)`` is the route handler itself (for API == server checks). + +Environment (read in the lifespan): ``HF_MODEL``, ``TT_MODEL_WEIGHTS_REVISION`` (then ``TT_WEIGHTS_REVISION``), +``_WEIGHTS_DIR``, ``TT_MESH_SHAPE`` (only 1x1), ``TT_DEVICE_ID``, ``_DISPATCH``, ``_NUM_CQS``, +``_VARIANT``, ``_WARMUP`` (JSON list, ``default`` or ``none``), ``_MAX_BODY_MB`` (default 256), +``_TORCH_THREADS``. +""" +import json +import logging +import math +import os +import re +import threading +import time +from contextlib import asynccontextmanager +from dataclasses import dataclass, field +from pathlib import Path +from typing import Any, Callable, Dict, List, Literal, Mapping, Optional, Type + +from fastapi import FastAPI, HTTPException, Request +from fastapi.responses import JSONResponse +from pydantic import BaseModel, ConfigDict, Field + +from .. import io as tio + +__all__ = [ + "HARDWARE", "PointsSpec", "SweepSpec", "ImageSpec", "StreamSpec", "ArraysSpec", "PredictRequest", "ServerSpec", + "ServerState", "parse_mesh_shape", "config_from_env", "default_model_factory", "decode_request", "create_app", +] + +HARDWARE = "Tenstorrent Blackhole p150 (single chip) via tt-nn" +log = logging.getLogger("ttaw.server") + + +# ------------------------------------------------------------------------------------------- request schema + +class _Strict(BaseModel): + model_config = ConfigDict(extra="forbid") # a misspelled field is a 422, not a silently ignored option + + +class PointsSpec(_Strict): + format: Literal["bin", "npy", "npz", "pcd", "list"] = "bin" + data: Optional[str] = None # base64 of the file bytes (all formats except "list") + values: Optional[List[List[float]]] = None # "list" only + fields: Optional[List[str]] = None # column names (required meaning for "bin" / "list") + dtype: Literal["float32", "float64"] = "float32" + key: Optional[str] = None # "npz" only + frame_id: Optional[str] = None + + +class SweepSpec(_Strict): + points: PointsSpec + time_lag_s: float = 0.0 # age of this sweep relative to the current one + T_current_from_sweep: Optional[Any] = None # 4x4 / {translation, rotation_wxyz} / {x, y, z, roll, pitch, yaw} + + +class ImageSpec(_Strict): + camera: str + data: str # base64 PNG / JPEG (or .npy uint8 HxWx3) + format: Literal["auto", "png", "jpeg", "npy"] = "auto" + intrinsics: Optional[Any] = None # 3x3 / 3x4 / {fx, fy, cx, cy} + T_ref_from_camera: Optional[Any] = None + distortion: Optional[Dict[str, Any]] = None + timestamp_s: Optional[float] = None + + +class StreamSpec(_Strict): + id: str = "default" + reset: bool = False + timestamp_s: Optional[float] = None + T_world_from_ego: Optional[Any] = None + + +class ArraysSpec(_Strict): + format: Literal["npz", "json"] = "npz" + data: Optional[str] = None + arrays: Optional[Dict[str, Any]] = None + + +class PredictRequest(_Strict): + """The common ``POST /predict`` envelope. Bundles with extra inputs subclass it (e.g. PointPainting ``rois``) + and decode the extra fields in ``ServerSpec.decode_extra``.""" + + points: Optional[PointsSpec] = None + sweeps: Optional[List[SweepSpec]] = None + images: Optional[List[ImageSpec]] = None + calibration: Optional[Dict[str, Any]] = None + stream: Optional[StreamSpec] = None + inputs: Optional[ArraysSpec] = None + params: Dict[str, Any] = Field(default_factory=dict) + output_format: Literal["json", "npz"] = "json" + + +# ----------------------------------------------------------------------------------------------- config + +@dataclass +class ServerSpec: + """Identity and hooks of one bundle's server. ``model_cls`` is the bundle's ``ModelBase`` subclass (used by the + default factory and for ``CAMERA_ORDER`` / calibration rules).""" + + model_name: str + env_prefix: str + model_cls: Any = None + task: str = "" + default_weights: str = "" + owner: str = "changh95" + hardware: str = HARDWARE + io: str = "" + autoware: Mapping[str, Any] = field(default_factory=dict) + source: Mapping[str, Any] = field(default_factory=dict) + calib_dir: Optional[Path] = None + default_variant: Optional[str] = None + version: str = "0.1.0" + description: str = "" + request_model: Type[PredictRequest] = PredictRequest + # decode_extra(request, kwargs, max_bytes): add model-specific call kwargs from extra request fields + decode_extra: Optional[Callable[[Any, Dict[str, Any], int], None]] = None + # info_extra(model) -> dict merged into /info + info_extra: Optional[Callable[[Any], Dict[str, Any]]] = None + + +_SHAPE_RE = re.compile(r"^\s*[\(\[]?\s*(\d+)\s*[xX,]\s*(\d+)\s*[\)\]]?\s*$") + + +def parse_mesh_shape(value: Optional[str]) -> tuple: + """``"1x1"`` / ``"(1, 1)"`` / ``"1,1"`` -> ``(1, 1)``; unset -> ``(1, 1)``.""" + if value is None or not value.strip(): + return (1, 1) + m = _SHAPE_RE.match(value) + if not m: + raise RuntimeError(f"TT_MESH_SHAPE={value!r} is not a mesh shape; expected '1x1', '(1, 1)' or '1,1'") + return (int(m.group(1)), int(m.group(2))) + + +def config_from_env(spec: ServerSpec, env: Optional[Mapping[str, str]] = None) -> Dict[str, Any]: + """The server configuration from the environment (read once, in the lifespan).""" + env = os.environ if env is None else env + p = spec.env_prefix + warm = (env.get(f"{p}_WARMUP") or "default").strip() or "default" + default_variant = spec.default_variant or getattr(spec.model_cls, "DEFAULT_VARIANT", None) or "default" + return { + "model_id": env.get("HF_MODEL") or spec.default_weights, + "revision": env.get("TT_MODEL_WEIGHTS_REVISION") or env.get("TT_WEIGHTS_REVISION") or None, + "weights_dir": env.get(f"{p}_WEIGHTS_DIR") or None, + "mesh_shape": parse_mesh_shape(env.get("TT_MESH_SHAPE")), + "device_id": int(env.get("TT_DEVICE_ID") or 0), + "dispatch": (env.get(f"{p}_DISPATCH") or "eth").strip().lower(), + "num_command_queues": int(env.get(f"{p}_NUM_CQS") or 0) or None, + "variant": env.get(f"{p}_VARIANT") or default_variant, + "warmup_variants": warm if warm in ("default", "none") else json.loads(warm), + "max_body_bytes": int(float(env.get(f"{p}_MAX_BODY_MB") or 256) * (1 << 20)), + "torch_threads": env.get(f"{p}_TORCH_THREADS") or None, + } + + +def default_model_factory(spec: ServerSpec) -> Callable[[Dict[str, Any]], Any]: + """``cfg -> spec.model_cls.from_pretrained(...)``.""" + + def factory(cfg: Dict[str, Any]) -> Any: + if spec.model_cls is None: + raise RuntimeError(f"{spec.model_name}: ServerSpec.model_cls is not set and no model_factory was given") + return spec.model_cls.from_pretrained(cfg["model_id"] or None, revision=cfg["revision"], + weights_dir=cfg["weights_dir"], variant=cfg["variant"], + device_id=cfg["device_id"], dispatch=cfg["dispatch"], + num_command_queues=cfg["num_command_queues"], + warmup_variants=cfg["warmup_variants"]) + + return factory + + +@dataclass +class ServerState: + """Per-app state (``app.state.ttaw``).""" + + spec: ServerSpec + model_factory: Callable[[Dict[str, Any]], Any] + ready: bool = False + error: Optional[str] = None + model: Any = None + config: Dict[str, Any] = field(default_factory=dict) + boot_ms: Optional[float] = None + lock: Any = field(default_factory=threading.Lock) # one chip, batch 1: device calls are serialized + predict: Optional[Callable[[Any], Dict[str, Any]]] = None + + def status(self) -> str: + return "ok" if self.ready else ("error" if self.error else "starting") + + def device_desc(self) -> Optional[Dict[str, Any]]: + return getattr(self.model, "device_info", None) if self.model is not None else None + + +def _non_finite_path(obj: Any, path: str = "body") -> Optional[str]: + """Location of the first NaN / infinite float in a JSON-able body (None if there is none).""" + if isinstance(obj, float): + return None if math.isfinite(obj) else path + if isinstance(obj, Mapping): + for k, v in obj.items(): + found = _non_finite_path(v, f"{path}.{k}") + if found is not None: + return found + elif isinstance(obj, (list, tuple)): + for i, v in enumerate(obj): + found = _non_finite_path(v, f"{path}[{i}]") + if found is not None: + return found + return None + + +def decode_request(req: PredictRequest, spec: ServerSpec, max_bytes: int) -> Dict[str, Any]: + """JSON envelope -> the Python-API call kwargs (all client errors raise ``io.InputError``).""" + cls = spec.model_cls + fields = tuple(getattr(cls, "POINT_FIELDS", tio.DEFAULT_POINT_FIELDS)) + kw: Dict[str, Any] = {} + if req.points is not None: + kw["points"] = tio.decode_points(req.points.model_dump(exclude_none=True), max_bytes=max_bytes, + default_fields=fields) + if req.sweeps: + kw["sweeps"] = [{"points": tio.decode_points(s.points.model_dump(exclude_none=True), max_bytes=max_bytes, + default_fields=fields), + "time_lag_s": s.time_lag_s, + "T_current_from_sweep": None if s.T_current_from_sweep is None else + tio.parse_transform(s.T_current_from_sweep, field="sweeps[].T_current_from_sweep")} + for s in req.sweeps] + calib = tio.resolve_calibration(req.calibration, spec.calib_dir) + if req.images: + order = tuple(getattr(cls, "CAMERA_ORDER", ()) or ()) or None + require = cls.requires_calibration() if hasattr(cls, "requires_calibration") else False + kw["images"] = tio.decode_cameras([i.model_dump(exclude_none=True) for i in req.images], calib, order=order, + require_calibration=require, max_bytes=max_bytes) + if calib is not None: + kw["calibration"] = calib + if req.stream is not None: + s = req.stream.model_dump(exclude_none=True) + if "T_world_from_ego" in s: + s["T_world_from_ego"] = tio.parse_transform(s["T_world_from_ego"], field="stream.T_world_from_ego") + kw["stream"] = s + if req.inputs is not None: + kw["inputs"] = tio.decode_named_arrays(req.inputs.model_dump(exclude_none=True), max_bytes=max_bytes) + if spec.decode_extra is not None: + spec.decode_extra(req, kw, max_bytes) + return kw + + +def create_app(spec: ServerSpec, *, model_factory: Optional[Callable[[Dict[str, Any]], Any]] = None) -> FastAPI: + """The ASGI app of one bundle (see the module docstring).""" + state = ServerState(spec, model_factory or default_model_factory(spec)) + RequestModel = spec.request_model + + @asynccontextmanager + async def lifespan(_app: FastAPI): + logging.basicConfig(level=logging.INFO, format="%(asctime)s %(levelname)s %(name)s: %(message)s") + cfg = config_from_env(spec) + if cfg["mesh_shape"] != (1, 1): + raise RuntimeError(f"TT_MESH_SHAPE={cfg['mesh_shape']} is a multi-chip mesh; {spec.model_name} runs on " + "exactly one Blackhole chip (mesh_device P150, shape 1x1)") + if cfg["torch_threads"]: + import torch + + torch.set_num_threads(int(cfg["torch_threads"])) + state.config, state.ready, state.error = cfg, False, None + t0 = time.perf_counter() + try: + state.model = state.model_factory(cfg) # logs: Loading weights / Opening device / Warming up / complete + except Exception as e: # startup failures exit uvicorn non-zero: no silent CPU fallback + state.error = f"{type(e).__name__}: {e}" + log.exception("startup failed") + raise + state.boot_ms = (time.perf_counter() - t0) * 1e3 + state.ready = True + log.info("%s ready in %.1f s (%s)", spec.model_name, state.boot_ms / 1e3, state.device_desc()) + try: + yield + finally: + state.ready = False + model, state.model = state.model, None + if model is not None: + log.info("Closing the model and the device") + with state.lock: + model.close() + + app = FastAPI(title=spec.model_name, version=spec.version, lifespan=lifespan, + description=spec.description or f"{spec.task} on one Tenstorrent Blackhole p150. " + "Not an OpenAI-compatible API.") + app.state.ttaw = state + + @app.middleware("http") + async def _body_limit(request: Request, call_next): + limit = (state.config or {}).get("max_body_bytes") + length = request.headers.get("content-length") + if limit and length and length.isdigit() and int(length) > limit * 4 // 3 + 4096: + return JSONResponse(status_code=400, content={"detail": f"request body of {length} bytes exceeds " + f"{spec.env_prefix}_MAX_BODY_MB"}) + return await call_next(request) + + @app.get("/") + def root() -> dict: + return {"model": spec.model_name, "routes": ["/health", "/v1/health", "/info", "/v1/models", "/predict"], + "docs": "/docs"} + + @app.get("/health") + def health() -> dict: + """Always 200: ``ok`` after warm-up, ``starting`` before, ``error`` if the boot failed.""" + return {"status": state.status(), "model": spec.model_name, "device": state.device_desc(), + "error": state.error} + + @app.get("/v1/health") + def v1_health() -> dict: # tt-model's hint for non-chat packages points here + return health() + + @app.get("/v1/models") + def v1_models() -> dict: + """OpenAI-shaped stub so tt-model's ready card / probes do not 404. This is NOT a chat API.""" + return {"object": "list", "data": [{"id": state.config.get("model_id") or spec.default_weights, + "object": "model", "owned_by": spec.owner}]} + + @app.get("/info") + def info() -> dict: + m = state.model + minfo = dict(getattr(m, "info", {}) or {}) if m is not None else {} + calib = sorted(p.stem for p in Path(spec.calib_dir).glob("*.json")) if spec.calib_dir else [] + body = { + "model": spec.model_name, "task": spec.task, "status": state.status(), "hardware": spec.hardware, + "io": spec.io, "autoware": dict(spec.autoware), + "weights": minfo.get("weights") or {"repo": state.config.get("model_id") or spec.default_weights}, + "device": state.device_desc(), "mesh_shape": os.environ.get("TT_MESH_SHAPE", "1x1"), + "input": {"kind": minfo.get("input_kind"), "point_fields": minfo.get("point_fields"), + "camera_order": minfo.get("camera_order"), "batch_size": 1, + "point_formats": list(tio.POINT_FORMATS), "max_body_bytes": state.config.get("max_body_bytes"), + "calibration_presets": calib, "extra_inputs": minfo.get("extra_inputs"), + "input_schema": minfo.get("input_schema")}, + "output": {"labels": minfo.get("labels")}, + "variant": minfo.get("variant"), "warm_variants": minfo.get("warm_variants"), + "runtime_params": minfo.get("runtime_params"), "compile_params": minfo.get("compile_params"), + "warmup_ms": minfo.get("warmup_ms"), "boot_ms": state.boot_ms, "source": dict(spec.source), + } + if spec.info_extra is not None and m is not None: + body.update(spec.info_extra(m)) + return body + + @app.post("/predict") + def predict(req: RequestModel) -> dict: + """One frame -> the model output's ``to_dict`` + ``timing_ms``. Sync handler: uvicorn runs it in a worker + thread; the lock serializes the device.""" + if not state.ready: + raise HTTPException(status_code=503, detail="model is still starting" if not state.error + else f"model failed to start: {state.error}") + model = state.model + t0 = time.perf_counter() + try: + kw = decode_request(req, spec, state.config.get("max_body_bytes") or (256 << 20)) + params = model.validate_params(req.params) + except tio.InputError as e: + raise HTTPException(status_code=400, detail=str(e)) from None + t1 = time.perf_counter() + try: + with state.lock: + out = model(**kw, **params) + except tio.InputError as e: # raised by the model's own pre-processing (shape / range checks) + raise HTTPException(status_code=400, detail=str(e)) from None + except Exception as e: # noqa: BLE001 -- a device / runtime failure + log.exception("inference failed") + raise HTTPException(status_code=500, detail=f"inference failed: {type(e).__name__}: {e}") from None + t2 = time.perf_counter() + body = out.to_dict(req.output_format) + bad = _non_finite_path(body) + if bad is not None: # strict JSON has no NaN / Infinity: say where, instead of a bare 500 from the encoder + log.error("inference produced a non-finite value at %s", bad) + raise HTTPException(status_code=500, detail=f"inference failed: non-finite value at {bad}") + timing = dict(body.get("timing_ms") or {}) + timing.update({"decode": round((t1 - t0) * 1e3, 3), "model_call": round((t2 - t1) * 1e3, 3), + "total": round((time.perf_counter() - t0) * 1e3, 3)}) + body["timing_ms"] = timing + return body + + state.predict = predict + return app diff --git a/code/tt_diffusion_planner/ttaw/server/client.py b/code/tt_diffusion_planner/ttaw/server/client.py new file mode 100644 index 0000000000000000000000000000000000000000..dbea7d078fcbe83c23ab89c714b62ede9eb19e92 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/server/client.py @@ -0,0 +1,186 @@ +#!/usr/bin/env python3 +# SPDX-License-Identifier: Apache-2.0 +"""C08 client: build / send ``POST /predict`` requests and run smoke checks. Standard library only. + +This module imports nothing from ttaw (no relative imports), so it works three ways: as ``.ttaw.server.client``, +as a script, and loaded by path from a bundle's ``smoke_test.py`` running under the host's system ``python3`` +(``container_smoke.sh``):: + + HERE = Path(__file__).resolve().parent # code//server + sys.path.insert(0, str(HERE.parent / "ttaw" / "server")) + import client as ttaw_client + +CLI:: + + python client.py --points sample.npz --out req.json # write the request + python client.py --points sample.npz --url http://127.0.0.1:20000 # send it, print the response + python client.py --image CAM_FRONT=f.jpg --calib-preset rig --url ... # cameras + python client.py --inputs scene.npz --url ... # planner +""" +from __future__ import annotations + +import argparse +import base64 +import json +import sys +import time +import urllib.error +import urllib.request +from pathlib import Path +from typing import Any, Dict, Iterable, List, Optional, Sequence, Tuple + +__all__ = ["POINT_SUFFIXES", "b64file", "build_request", "post", "get_json", "wait_ready", "check_service", + "check_bad_request", "smoke_line", "main"] + +POINT_SUFFIXES = {".bin": "bin", ".npy": "npy", ".npz": "npz", ".pcd": "pcd"} + + +def b64file(path: Any) -> str: + return base64.b64encode(Path(path).read_bytes()).decode("ascii") + + +def build_request(points: Optional[str] = None, fields: Optional[Sequence[str]] = None, images: Iterable[str] = (), + calib: Optional[str] = None, calib_preset: Optional[str] = None, inputs: Optional[str] = None, + params: Optional[Dict[str, Any]] = None, stream_id: Optional[str] = None, reset: bool = False, + output_format: str = "json") -> Dict[str, Any]: + """A ``/predict`` body from files: ``points`` path (.bin needs ``fields``), ``images`` as ``NAME=PATH``, + ``calib`` JSON file or ``calib_preset`` name, ``inputs`` .npz (planner), ``params``, stream id / reset.""" + req: Dict[str, Any] = {} + if points: + p = Path(points) + fmt = POINT_SUFFIXES.get(p.suffix.lower()) + if fmt is None: + raise ValueError(f"points: unknown suffix {p.suffix!r} (use .bin / .npy / .npz / .pcd)") + req["points"] = {"format": fmt, "data": b64file(p)} + if fields: + req["points"]["fields"] = list(fields) + images = list(images) + if images: + req["images"] = [] + for spec in images: + name, _, path = spec.partition("=") + if not path: + raise ValueError(f"image expects NAME=PATH, got {spec!r}") + req["images"].append({"camera": name, "data": b64file(path)}) + if calib: + req["calibration"] = json.loads(Path(calib).read_text()) + elif calib_preset: + req["calibration"] = {"preset": calib_preset} + if inputs: + req["inputs"] = {"format": "npz", "data": b64file(inputs)} + if params: + req["params"] = dict(params) + if stream_id is not None or reset: + req["stream"] = {"id": stream_id or "default", "reset": bool(reset)} + if output_format != "json": + req["output_format"] = output_format + return req + + +def post(url: str, payload: Dict[str, Any], timeout: float = 120.0) -> Tuple[int, Dict[str, Any]]: + """POST ``payload`` to ``/predict`` -> ``(status, json body)`` (HTTP errors are returned, not raised).""" + data = json.dumps(payload).encode("utf-8") + r = urllib.request.Request(url.rstrip("/") + "/predict", data=data, headers={"Content-Type": "application/json"}) + try: + with urllib.request.urlopen(r, timeout=timeout) as resp: + return resp.status, json.loads(resp.read().decode("utf-8")) + except urllib.error.HTTPError as e: + body = e.read().decode("utf-8", "replace") + try: + return e.code, json.loads(body) + except ValueError: + return e.code, {"detail": body} + + +def get_json(url: str, timeout: float = 10.0) -> Dict[str, Any]: + with urllib.request.urlopen(url, timeout=timeout) as r: + return json.loads(r.read().decode("utf-8")) + + +def wait_ready(base: str, wait_s: float = 0.0, poll_s: float = 5.0) -> Dict[str, Any]: + """Poll ``/health`` until ``ok`` / ``error`` or ``wait_s`` elapses; returns the last health body.""" + deadline = time.time() + wait_s + last: Dict[str, Any] = {} + while True: + try: + last = get_json(base.rstrip("/") + "/health") + if last.get("status") in ("ok", "error"): + return last + except (urllib.error.URLError, OSError, ValueError) as e: + last = {"status": f"unreachable ({e})"} + if time.time() >= deadline: + return last + time.sleep(poll_s) + + +def check_service(base: str, *, expect_dispatch: Optional[str] = "eth", + expect_grid: Optional[str] = "12x10") -> Tuple[Dict[str, Any], List[str]]: + """``/info`` and ``/v1/models`` answer; ``/info`` reports the pinned dispatch and grid (PLAN.md 0.3 item 6: + the smoke must assert ETH dispatch and 12x10, not just print them). Returns ``(info, failures)``.""" + fails: List[str] = [] + try: + info = get_json(base.rstrip("/") + "/info") + models = get_json(base.rstrip("/") + "/v1/models") + except (urllib.error.URLError, OSError, ValueError) as e: + return {}, [f"/info or /v1/models: {e}"] + if not models.get("data"): + fails.append("/v1/models is empty") + device = info.get("device") or {} + if expect_dispatch and device.get("dispatch") != expect_dispatch: + fails.append(f"dispatch is {device.get('dispatch')!r}, expected {expect_dispatch!r}") + if expect_grid and device.get("grid") != expect_grid: + fails.append(f"grid is {device.get('grid')!r}, expected {expect_grid!r}") + return info, fails + + +def check_bad_request(base: str, payload: Optional[Dict[str, Any]] = None) -> Optional[str]: + """A malformed request must answer 400 (not 500). Returns a failure message or None.""" + code, _ = post(base, payload or {"points": {"format": "npy", "data": "not base64!"}}) + return None if code == 400 else f"malformed request answered {code}, expected 400" + + +def smoke_line(model: str, summary: str, failures: Sequence[str]) -> str: + """The one-line verdict container_smoke.sh greps: ``PASS : ...`` / ``FAIL : ... :: reasons``.""" + if failures: + return f"FAIL {model}: {summary} :: " + "; ".join(list(failures)[:5]) + return f"PASS {model}: {summary}" + + +def _param(text: str) -> Tuple[str, Any]: + k, _, v = text.partition("=") + try: + return k, json.loads(v) + except ValueError: + return k, v + + +def main(argv: Optional[Sequence[str]] = None) -> int: + ap = argparse.ArgumentParser(description=__doc__, formatter_class=argparse.RawDescriptionHelpFormatter) + ap.add_argument("--points") + ap.add_argument("--fields", help="comma list, e.g. x,y,z,intensity (needed for .bin)") + ap.add_argument("--image", action="append", default=[], help="NAME=PATH, repeat per camera") + ap.add_argument("--calib", help="calibration JSON file") + ap.add_argument("--calib-preset", help="a preset name listed by /info") + ap.add_argument("--inputs", help=".npz of named input tensors (planner)") + ap.add_argument("--param", action="append", default=[], help="k=v (JSON value), repeat") + ap.add_argument("--stream-id") + ap.add_argument("--reset", action="store_true") + ap.add_argument("--output-format", default="json", choices=["json", "npz"]) + ap.add_argument("--out", help="write the request JSON here") + ap.add_argument("--url", help="POST it to this server and print the response") + a = ap.parse_args(argv) + req = build_request(a.points, a.fields.split(",") if a.fields else None, a.image, a.calib, a.calib_preset, + a.inputs, dict(_param(p) for p in a.param), a.stream_id, a.reset, a.output_format) + if a.out: + Path(a.out).write_text(json.dumps(req)) + if a.url: + code, body = post(a.url, req) + print(json.dumps(body, indent=1)[:20000]) + return 0 if code == 200 else 1 + if not a.out: + json.dump(req, sys.stdout) + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/code/tt_diffusion_planner/ttaw/server/smoke.py b/code/tt_diffusion_planner/ttaw/server/smoke.py new file mode 100644 index 0000000000000000000000000000000000000000..ba75158cec2288b7f9ed9555a9d206da625bee6b --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/server/smoke.py @@ -0,0 +1,482 @@ +#!/usr/bin/env python3 +# SPDX-License-Identifier: Apache-2.0 +"""C08 smoke checks beyond HTTP: does a served ``/predict`` body agree with the stored CPU-reference body of the +same sample, and does ``/info`` report what the staged tt-model package pins for the served profile. Standard +library only. + +Like :mod:`client`, this module imports nothing from ttaw, so a bundle's ``smoke_test.py`` (run by +``container_smoke.sh`` with the host's ``python3``, which has neither numpy nor the bundle) loads it by path:: + + sys.path.insert(0, str(HERE.parent / "ttaw" / "server")) # HERE = code//server + import client as ttaw_client + import smoke as ttaw_smoke + + info, fails = ttaw_client.check_service(url, expect_dispatch="eth", expect_grid="12x10") + pinned = ttaw_smoke.pinned_config(f"{staged}/tt_kernel_manifest.json", profile) + fails += ttaw_smoke.check_pinned(info, pinned, "CENTERPOINT") # /info runs what serve.env pins + ref = ttaw_smoke.find_reference(SAMPLE, pinned["profile"], info.get("variant")) + metrics, more = ttaw_smoke.compare_with_reference(body, ttaw_smoke.load_json(ref), gates={"min_recall": 0.9}) + +The stored reference is the ``to_dict()`` body of the fp32 CPU reference on the shipped sample, saved next to it as +``.reference.json``, or ``..reference.json`` / ``..reference.json`` +when the output depends on the serve profile or the model variant. Synthetic samples may give no detections, so +agreement with the reference, not a detection count, is their smoke gate (PLAN.md section 6.3). + +What is compared, by the keys of the reference body (BUNDLE_CONVENTIONS.md section 7.4): + +- ``detections``: greedy same-label matching, reference rows in descending score order; 3-D rows (``center``) + match by BEV centre distance, 2-D rows (``box_xyxy``) by IoU; gates on recall, precision and the largest score + difference of a matched pair; +- ``trajectory``: average / final displacement over the ``x`` and ``y`` columns; +- every encoded array (``ttaw.io.encode_array`` npz / npz_compressed / raw / list, ``encode_png``), e.g. + per-point ``labels``, a ``mask`` or a ``semseg`` extra: integer arrays by the fraction of equal elements, float + arrays by the largest absolute difference; +- ``model`` and ``frame_id`` must be equal. +""" +from __future__ import annotations + +import ast +import base64 +import io +import json +import math +import re +import struct +import sys +import zipfile +import zlib +from pathlib import Path +from typing import Any, Callable, Dict, List, Mapping, NamedTuple, Optional, Sequence, Tuple, Union + +__all__ = ["DEFAULT_GATES", "ARRAY_FORMATS", "Array", "load_json", "decode_array", "find_reference", + "compare_with_reference", "describe_metrics", "pinned_config", "check_pinned"] + +ARRAY_FORMATS = ("npz", "npz_compressed", "raw", "list", "png") +DEFAULT_GATES: Dict[str, Optional[float]] = { + "max_center_dist": 0.5, # 3-D detections: same-label BEV centre distance of a match (m) + "min_iou": 0.5, # 2-D detections: same-label box IoU of a match + "min_recall": 0.95, # matched / reference detections + "min_precision": 0.95, # matched / served detections + "max_score_err": 0.05, # largest |score difference| of a matched pair + "min_label_agreement": 0.99, # integer arrays (per-point labels, masks): fraction of equal elements + "max_abs_err": None, # float arrays: largest |difference| (reported only, unless set) + "max_ade": 0.5, # trajectories: average displacement over x, y (m) + "max_fde": 1.0, # trajectories: final displacement (m) +} + +PathLike = Union[str, Path] +# numpy type code (without the byte-order character) -> (struct code, numpy dtype name) +_TYPES = {"b1": ("?", "bool"), "u1": ("B", "uint8"), "u2": ("H", "uint16"), "u4": ("I", "uint32"), + "u8": ("Q", "uint64"), "i1": ("b", "int8"), "i2": ("h", "int16"), "i4": ("i", "int32"), + "i8": ("q", "int64"), "f2": ("e", "float16"), "f4": ("f", "float32"), "f8": ("d", "float64")} +_BY_NAME = {name: code for code, name in _TYPES.values()} +_SAFE_KEY = re.compile(r"[A-Za-z0-9_.-]+") + + +class Array(NamedTuple): + """A decoded array: ``values`` are the elements in C order (flattened).""" + + shape: Tuple[int, ...] + dtype: str + values: List[Any] + + @property + def is_float(self) -> bool: + return self.dtype.startswith("float") + + +def load_json(path: PathLike) -> Dict[str, Any]: + return json.loads(Path(path).read_text()) + + +# ------------------------------------------------------------------------------------------------ decoding + +def _count(shape: Sequence[int]) -> int: + n = 1 + for d in shape: + n *= int(d) + return n + + +def _fortran_to_c(values: List[Any], shape: Sequence[int]) -> List[Any]: + """Reorder column-major elements into row-major order.""" + strides, step = [], 1 + for d in shape: + strides.append(step) + step *= d + out, index = [], [0] * len(shape) + for _ in range(len(values)): + out.append(values[sum(i * s for i, s in zip(index, strides))]) + for k in range(len(shape) - 1, -1, -1): + index[k] += 1 + if index[k] < shape[k]: + break + index[k] = 0 + return out + + +def _npy(raw: bytes) -> Array: + """One ``.npy`` file (format 1.0 / 2.0 / 3.0, plain numeric dtypes).""" + if raw[:6] != b"\x93NUMPY": + raise ValueError("not an .npy file") + major = raw[6] + if major == 1: + (hlen,), start = struct.unpack("": ">", "=": "<" if sys.byteorder == "little" else ">"}.get(descr[0], "<") + code, name = _TYPES[descr[1:]] + n = _count(shape) + data = raw[start + hlen:] + size = n * struct.calcsize("<" + code) + if len(data) < size: + raise ValueError(f".npy data holds {len(data)} bytes, {size} expected") + values = list(struct.unpack(f"{order}{n}{code}", data[:size])) + if header.get("fortran_order") and len(shape) > 1: + values = _fortran_to_c(values, shape) + return Array(shape, name, values) + + +def _paeth(a: int, b: int, c: int) -> int: + p = a + b - c + pa, pb, pc = abs(p - a), abs(p - b), abs(p - c) + if pa <= pb and pa <= pc: + return a + return b if pb <= pc else c + + +def _png(raw: bytes) -> Array: + """An 8-bit, non-interlaced grey / grey+alpha / RGB / RGBA PNG (what ``ttaw.io.encode_png`` writes).""" + if raw[:8] != b"\x89PNG\r\n\x1a\n": + raise ValueError("not a PNG file") + pos, idat, head = 8, [], None + while pos + 8 <= len(raw): + length, kind = struct.unpack(">I4s", raw[pos:pos + 8]) + chunk = raw[pos + 8:pos + 8 + length] + pos += 12 + length + if kind == b"IHDR": + head = struct.unpack(">IIBBBBB", chunk) + elif kind == b"IDAT": + idat.append(chunk) + elif kind == b"IEND": + break + if head is None: + raise ValueError("PNG without IHDR") + width, height, depth, color, _, _, interlace = head + channels = {0: 1, 2: 3, 4: 2, 6: 4}.get(color) + if depth != 8 or channels is None or interlace: + raise ValueError(f"PNG bit depth {depth}, colour type {color}, interlace {interlace} is not supported") + data = zlib.decompress(b"".join(idat)) + stride, bpp = width * channels, channels + if len(data) < height * (stride + 1): + raise ValueError("truncated PNG image data") + out, prev = bytearray(), bytearray(stride) + for y in range(height): + base = y * (stride + 1) + ftype, row = data[base], bytearray(data[base + 1:base + 1 + stride]) + if ftype == 1: + for i in range(bpp, stride): + row[i] = (row[i] + row[i - bpp]) & 0xFF + elif ftype == 2: + for i in range(stride): + row[i] = (row[i] + prev[i]) & 0xFF + elif ftype == 3: + for i in range(stride): + left = row[i - bpp] if i >= bpp else 0 + row[i] = (row[i] + ((left + prev[i]) >> 1)) & 0xFF + elif ftype == 4: + for i in range(stride): + left, upper_left = (row[i - bpp], prev[i - bpp]) if i >= bpp else (0, 0) + row[i] = (row[i] + _paeth(left, prev[i], upper_left)) & 0xFF + elif ftype != 0: + raise ValueError(f"PNG filter type {ftype} is invalid") + out += row + prev = row + shape = (height, width) if channels == 1 else (height, width, channels) + return Array(shape, "uint8", list(out)) + + +def _flatten(value: Any, out: List[Any]) -> List[Any]: + if isinstance(value, list): + for v in value: + _flatten(v, out) + else: + out.append(value) + return out + + +def decode_array(spec: Mapping[str, Any]) -> Array: + """Decode an encoded output array ``{"format", "key", "dtype", "shape", "data"}`` (``ttaw.io.encode_array`` / + ``encode_png``) without numpy. Raises ``ValueError`` for anything it cannot decode.""" + fmt = spec.get("format") + if fmt not in ARRAY_FORMATS: + raise ValueError(f"array format {fmt!r} is not one of {ARRAY_FORMATS}") + if fmt == "list": + dtype = str(spec.get("dtype") or "float64") + arr = Array(tuple(int(d) for d in spec.get("shape") or ()), dtype, _flatten(spec.get("data"), [])) + else: + raw = base64.b64decode(spec.get("data") or "") + if fmt == "png": + arr = _png(raw) + elif fmt == "raw": + dtype = str(spec.get("dtype")) + if dtype not in _BY_NAME: + raise ValueError(f"raw array dtype {dtype!r} is not supported") + shape = tuple(int(d) for d in spec.get("shape") or ()) + n, code = _count(shape), _BY_NAME[dtype] + if len(raw) != n * struct.calcsize("<" + code): + raise ValueError(f"raw array holds {len(raw)} bytes, shape {list(shape)} {dtype} needs " + f"{n * struct.calcsize('<' + code)}") + arr = Array(shape, dtype, list(struct.unpack(f"<{n}{code}", raw))) + else: + with zipfile.ZipFile(io.BytesIO(raw)) as z: + names = z.namelist() + member = f"{spec.get('key')}.npy" + if member not in names: + if len(names) != 1: + raise ValueError(f"npz has no member {member!r} (members {names})") + member = names[0] + arr = _npy(z.read(member)) + if "shape" in spec and list(arr.shape) != [int(d) for d in spec["shape"]]: + raise ValueError(f"decoded shape {list(arr.shape)} differs from the declared {list(spec['shape'])}") + if len(arr.values) != _count(arr.shape): + raise ValueError(f"{len(arr.values)} elements do not fill shape {list(arr.shape)}") + return arr + + +def _is_encoded_array(value: Any) -> bool: + return isinstance(value, Mapping) and value.get("format") in ARRAY_FORMATS and "data" in value + + +# ---------------------------------------------------------------------------------------------- comparison + +def find_reference(sample: PathLike, *keys: Optional[str]) -> Optional[Path]: + """The stored CPU-reference body of ``sample``: the first existing ``..reference.json`` for the + given keys (e.g. the serve profile, then the model variant), else ``.reference.json``; None if absent.""" + p = Path(sample) + names = [f"{p.stem}.{k}.reference.json" for k in keys if k and _SAFE_KEY.fullmatch(str(k))] + for name in names + [f"{p.stem}.reference.json"]: + if (p.parent / name).is_file(): + return p.parent / name + return None + + +def _label(row: Mapping[str, Any]) -> Any: + return row.get("label_id", row.get("label")) + + +def _iou(a: Sequence[float], b: Sequence[float]) -> float: + ix = max(0.0, min(a[2], b[2]) - max(a[0], b[0])) + iy = max(0.0, min(a[3], b[3]) - max(a[1], b[1])) + inter = ix * iy + union = (a[2] - a[0]) * (a[3] - a[1]) + (b[2] - b[0]) * (b[3] - b[1]) - inter + return inter / union if union > 0 else 0.0 + + +def _greedy_match(test: List[Mapping[str, Any]], ref: List[Mapping[str, Any]], + cost: Callable[[Mapping[str, Any], Mapping[str, Any]], float], + accept: Callable[[float], bool]) -> List[Tuple[int, int, float]]: + """Reference rows in descending score order (stable) each take the cheapest unused same-label test row; + returns ``(test index, reference index, cost)`` of the accepted pairs.""" + used = [False] * len(test) + pairs = [] + for i in sorted(range(len(ref)), key=lambda k: -float(ref[k].get("score", 0.0))): + best, best_cost = -1, math.inf + for j, row in enumerate(test): + if not used[j] and _label(row) == _label(ref[i]): + c = cost(row, ref[i]) + if c < best_cost: + best, best_cost = j, c + if best >= 0 and accept(best_cost): + used[best] = True + pairs.append((best, i, best_cost)) + return pairs + + +def _compare_detections(test: Any, ref: List[Mapping[str, Any]], + g: Mapping[str, Any]) -> Tuple[Dict[str, Any], List[str]]: + if not isinstance(test, list): + return {}, ["the response has no 'detections' list"] + rows = ref + test + if any("center" in d for d in rows): + kind = "3d" + pairs = _greedy_match(test, ref, lambda t, r: math.hypot(t["center"][0] - r["center"][0], + t["center"][1] - r["center"][1]), + lambda c: c <= g["max_center_dist"]) + elif any("box_xyxy" in d for d in rows): + kind = "2d" + pairs = _greedy_match(test, ref, lambda t, r: 1.0 - _iou(t["box_xyxy"], r["box_xyxy"]), + lambda c: 1.0 - c >= g["min_iou"]) + elif not rows: # nothing detected on either side (e.g. a synthetic sample): agreement + kind, pairs = "none", [] + else: + return {}, ["detections carry neither 'center' nor 'box_xyxy'"] + score_err = max((abs(float(test[j]["score"]) - float(ref[i]["score"])) for j, i, _ in pairs), default=0.0) + m: Dict[str, Any] = {"kind": kind, "n": len(test), "n_ref": len(ref), "matched": len(pairs), + "recall": len(pairs) / len(ref) if ref else 1.0, + "precision": len(pairs) / len(test) if test else 1.0, "score_err_max": score_err} + if kind == "3d": + m["center_err_max"] = max((c for _, _, c in pairs), default=0.0) + elif kind == "2d": + m["iou_min"] = min((1.0 - c for _, _, c in pairs), default=1.0) + fails = [] + if m["recall"] < g["min_recall"]: + fails.append(f"detections recall {m['recall']:.3f} < {g['min_recall']} ({len(pairs)}/{len(ref)} reference " + "detections matched)") + if m["precision"] < g["min_precision"]: + fails.append(f"detections precision {m['precision']:.3f} < {g['min_precision']} ({len(pairs)}/{len(test)} " + "served detections matched)") + if score_err > g["max_score_err"]: + fails.append(f"detections max |score difference| {score_err:.4f} > {g['max_score_err']}") + return m, fails + + +def _compare_trajectory(body: Mapping[str, Any], ref: Mapping[str, Any], + g: Mapping[str, Any]) -> Tuple[Dict[str, Any], List[str]]: + test, want = body.get("trajectory"), ref["trajectory"] + if not isinstance(test, list): + return {}, ["the response has no 'trajectory'"] + if len(test) != len(want): + return {"n": len(test), "n_ref": len(want)}, [f"trajectory has {len(test)} poses, the reference {len(want)}"] + columns = list(ref.get("columns") or ["x", "y"]) + if body.get("columns", columns) != columns: + return {}, [f"trajectory columns {body.get('columns')} differ from the reference {columns}"] + ix, iy = (columns.index("x"), columns.index("y")) if {"x", "y"} <= set(columns) else (0, 1) + disp = [math.hypot(a[ix] - b[ix], a[iy] - b[iy]) for a, b in zip(test, want)] + m = {"n": len(test), "ade": sum(disp) / len(disp) if disp else 0.0, "fde": disp[-1] if disp else 0.0} + fails = [] + if not m["ade"] <= g["max_ade"]: + fails.append(f"trajectory ADE {m['ade']:.4f} > {g['max_ade']}") + if not m["fde"] <= g["max_fde"]: + fails.append(f"trajectory FDE {m['fde']:.4f} > {g['max_fde']}") + return m, fails + + +def _compare_array(name: str, test: Any, ref: Mapping[str, Any], + g: Mapping[str, Any]) -> Tuple[Dict[str, Any], List[str]]: + if not _is_encoded_array(test): + return {}, [f"the response has no encoded array {name!r}"] + a, b = decode_array(test), decode_array(ref) + if a.shape != b.shape: + return {"shape": list(a.shape), "shape_ref": list(b.shape)}, [ + f"{name}: shape {list(a.shape)} differs from the reference {list(b.shape)}"] + if a.is_float or b.is_float: + err = 0.0 + for x, y in zip(a.values, b.values): + if x == y: # equal values, equal infinities included + continue + d = abs(float(x) - float(y)) + if d != d: # NaN on either side + err = math.nan + break + err = max(err, d) + limit = g["max_abs_err"] + fails = [f"{name}: max |difference| {err:.6g} > {limit}"] if limit is not None and not err <= limit else [] + return {"max_abs_err": err, "n": len(b.values)}, fails + equal = sum(1 for x, y in zip(a.values, b.values) if x == y) + agreement = equal / len(b.values) if b.values else 1.0 + fails = [] + if agreement < g["min_label_agreement"]: + fails.append(f"{name}: agreement {agreement:.4f} < {g['min_label_agreement']} " + f"({len(b.values) - equal} of {len(b.values)} elements differ)") + return {"agreement": agreement, "n": len(b.values)}, fails + + +def compare_with_reference(body: Mapping[str, Any], reference: Mapping[str, Any], + gates: Optional[Mapping[str, Optional[float]]] = None + ) -> Tuple[Dict[str, Dict[str, Any]], List[str]]: + """Compare a served ``/predict`` body with the stored CPU-reference body of the same input (module docstring). + ``gates`` override :data:`DEFAULT_GATES` (unknown names raise ``ValueError``). Returns ``(metrics per compared + key, failures)``; a body that cannot be decoded is a failure, not an exception.""" + unknown = sorted(set(gates or {}) - set(DEFAULT_GATES)) + if unknown: + raise ValueError(f"unknown reference gates {unknown}; known: {sorted(DEFAULT_GATES)}") + g = {**DEFAULT_GATES, **dict(gates or {})} + metrics: Dict[str, Dict[str, Any]] = {} + fails = [f"{k} is {body.get(k)!r}, the reference has {reference[k]!r}" for k in ("model", "frame_id") + if k in reference and body.get(k) != reference[k]] + checks: List[Tuple[str, Callable[[], Tuple[Dict[str, Any], List[str]]]]] = [] + if "detections" in reference: + checks.append(("detections", lambda: _compare_detections(body.get("detections"), reference["detections"], g))) + if "trajectory" in reference: + checks.append(("trajectory", lambda: _compare_trajectory(body, reference, g))) + for key, value in reference.items(): + if key != "arrays" and _is_encoded_array(value): + checks.append((key, lambda k=key, v=value: _compare_array(k, body.get(k), v, g))) + if not checks: + fails.append("the reference holds nothing to compare (no detections, trajectory or encoded arrays)") + for key, check in checks: + try: + metrics[key], more = check() + except (ValueError, KeyError, TypeError, IndexError, struct.error, zlib.error, zipfile.BadZipFile) as e: + metrics[key], more = {}, [f"{key}: cannot compare ({type(e).__name__}: {e})"] + fails += more + return metrics, fails + + +def describe_metrics(metrics: Mapping[str, Mapping[str, Any]]) -> str: + """One compact line for the smoke verdict, e.g. ``detections 5/5 recall=1.000 precision=1.000 dscore=0.0040``.""" + parts = [] + for key, m in metrics.items(): + if "recall" in m: + parts.append(f"{key} {m['matched']}/{m['n_ref']} recall={m['recall']:.3f} " + f"precision={m['precision']:.3f} dscore={m['score_err_max']:.4f}") + elif "ade" in m: + parts.append(f"{key} ade={m['ade']:.4f} fde={m['fde']:.4f}") + elif "agreement" in m: + parts.append(f"{key} agree={m['agreement']:.4f}") + elif "max_abs_err" in m: + parts.append(f"{key} max_abs={m['max_abs_err']:.4g}") + else: + parts.append(f"{key} -") + return "; ".join(parts) or "-" + + +# ----------------------------------------------------------------------------------------- package pins + +def pinned_config(manifest: PathLike, profile: Optional[str] = None) -> Dict[str, Any]: + """What a staged package (``//tt_kernel_manifest.json``) pins for one serve profile (default: the + package's default profile): ``{"profile": name, "env": serve.env with the profile's env on top (tt-model's + merge rule), "weights": {"repo", "revision"}}``.""" + m = load_json(manifest) + c = m.get("container") or {} + profiles = c.get("serve_profiles") or [{"name": "default", "env": {}}] + name = profile or c.get("default_profile") or profiles[0].get("name") + chosen = [p for p in profiles if p.get("name") == name] + if not chosen: + raise ValueError(f"no serve profile {name!r} in {manifest}; available: {[p.get('name') for p in profiles]}") + env = dict((c.get("serve") or {}).get("env") or {}) + env.update({k: v for k, v in (chosen[0].get("env") or {}).items() if v is not None}) + weights = m.get("weights") or {} + return {"profile": name, "env": {k: str(v) for k, v in env.items()}, + "weights": {"repo": weights.get("repo_id"), "revision": weights.get("revision")}} + + +def check_pinned(info: Mapping[str, Any], pinned: Mapping[str, Any], prefix: str) -> List[str]: + """``/info`` runs what the package pins: ``_DISPATCH`` (eth | worker), ``_NUM_CQS``, + ``_VARIANT`` and the weights revision. Pins that are absent are not checked. Returns failures.""" + env = pinned.get("env") or {} + device = info.get("device") or {} + profile = pinned.get("profile") + fails = [] + dispatch = (env.get(f"{prefix}_DISPATCH") or "").strip().lower() + if dispatch in ("eth", "worker") and device.get("dispatch") != dispatch: + fails.append(f"dispatch is {device.get('dispatch')!r}, profile {profile!r} pins {prefix}_DISPATCH={dispatch}") + cqs = (env.get(f"{prefix}_NUM_CQS") or "").strip() + if cqs and str(device.get("num_command_queues")) != cqs: + fails.append(f"num_command_queues is {device.get('num_command_queues')!r}, profile {profile!r} pins " + f"{prefix}_NUM_CQS={cqs}") + variant = (env.get(f"{prefix}_VARIANT") or "").strip() + if variant and info.get("variant") != variant: + fails.append(f"variant is {info.get('variant')!r}, profile {profile!r} pins {prefix}_VARIANT={variant}") + revision = (pinned.get("weights") or {}).get("revision") + served = (info.get("weights") or {}).get("revision") + if revision and served != revision: + fails.append(f"weights revision is {served!r}, the package pins {revision!r}") + return fails diff --git a/code/tt_diffusion_planner/ttaw/tensors.py b/code/tt_diffusion_planner/ttaw/tensors.py new file mode 100644 index 0000000000000000000000000000000000000000..34b69ea1da9e97a8851236edf35cf171f1cef0ff --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/tensors.py @@ -0,0 +1,363 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C03: tensor layout and dtype helpers shared by every port. + +Host side (numpy only, no ttnn): tile padding, fixed-capacity padding with a valid count, bf16 bit conversion. +Device side (ttnn imported inside each function): exact host-tensor construction for every dtype (integer arrays are +never routed through a float intermediate), uploads, layout and dtype changes that are no-ops when nothing changes, +fp32 islands (typecast in and out), uint8 uploads with an on-device typecast, readback to numpy, and +:class:`HostStaging`: a persistent host tensor whose buffer is written in place (no ``from_torch`` per frame). + +Measured on this p150b (``logs/ttaw/api_probe_20261007.log``): numpy uint32 arrays round-trip exactly (including +values >= 2**31); uint8 -> bf16 ``ttnn.typecast`` is exact in ROW_MAJOR and TILE layout; ``torch.from_dlpack`` of a +host buffer is a writable alias (``numpy.from_dlpack`` is read-only, ``numpy.asarray`` fails). +""" +from __future__ import annotations + +import math +from typing import Any, Callable, Optional, Sequence, Tuple, Union + +import numpy as np + +__all__ = [ + "TILE", + "round_up", + "tile_padded_shape", + "pad_to_tile", + "pad_to_capacity", + "as_4d", + "float32_to_bf16_bits", + "bf16_bits_to_float32", + "round_to_bf16", + "ttnn_dtype", + "dtype_name", + "to_host_tensor", + "to_device", + "to_numpy", + "to_layout", + "typecast", + "to_fp32", + "to_bf16", + "fp32_island", + "upload_u8", + "HostStaging", +] + +TILE = 32 + +ArrayLike = Union[np.ndarray, Sequence, Any] + + +# ------------------------------------------------------------------------------------------------- host side + +def round_up(n: int, multiple: int = TILE) -> int: + """Smallest multiple of ``multiple`` that is >= ``n``.""" + if multiple <= 0: + raise ValueError("multiple must be positive") + return -(-int(n) // multiple) * multiple + + +def tile_padded_shape(shape: Sequence[int]) -> Tuple[int, ...]: + """``shape`` with its last two dims rounded up to the 32x32 tile (what a TILE tensor occupies).""" + shape = tuple(int(s) for s in shape) + if len(shape) < 2: + raise ValueError(f"need at least 2 dims, got {shape}") + return shape[:-2] + (round_up(shape[-2]), round_up(shape[-1])) + + +def pad_to_tile(array: np.ndarray, value: float = 0.0) -> np.ndarray: + """Pad the last two dims of a host array up to multiples of 32 with ``value`` (copy only when needed).""" + a = np.asarray(array) + target = tile_padded_shape(a.shape) + if target == a.shape: + return a + pad = [(0, 0)] * (a.ndim - 2) + [(0, target[-2] - a.shape[-2]), (0, target[-1] - a.shape[-1])] + return np.pad(a, pad, mode="constant", constant_values=value) + + +def pad_to_capacity(array: np.ndarray, capacity: int, *, axis: int = 0, value: float = 0.0, + truncate: bool = False) -> Tuple[np.ndarray, int]: + """Fixed-capacity buffer for a traced graph: pad ``axis`` to ``capacity`` rows with ``value``. + + Returns ``(padded, n_valid)``. More than ``capacity`` rows raise ``ValueError`` unless ``truncate`` keeps the + first ``capacity`` (the caller must report the overflow). Capacities are grid-independent constants + (PLAN.md section 0.2).""" + a = np.asarray(array) + n = a.shape[axis] + if n > capacity: + if not truncate: + raise ValueError(f"{n} rows exceed the capacity {capacity}") + return np.take(a, np.arange(capacity), axis=axis), capacity + if n == capacity: + return a, n + pad = [(0, 0)] * a.ndim + pad[axis] = (0, capacity - n) + return np.pad(a, pad, mode="constant", constant_values=value), n + + +def as_4d(array: np.ndarray) -> np.ndarray: + """Prepend unit dims up to rank 4 (the ``[N, C, H, W]`` / ``[1, 1, rows, cols]`` convention of ttnn).""" + a = np.asarray(array) + if a.ndim > 4: + raise ValueError(f"rank {a.ndim} > 4") + return a.reshape((1,) * (4 - a.ndim) + a.shape) + + +def float32_to_bf16_bits(array: ArrayLike) -> np.ndarray: + """float32 -> bfloat16 bit patterns (uint16), round-to-nearest-even like ``torch.Tensor.to(torch.bfloat16)``. + NaN stays NaN (quiet bit forced).""" + f = np.ascontiguousarray(array, dtype=np.float32) + bits = f.view(np.uint32).astype(np.uint64) + rounded = (bits + 0x7FFF + ((bits >> 16) & 1)) >> 16 + out = rounded.astype(np.uint16) + nan = np.isnan(f) + if nan.any(): + out[nan] = ((bits[nan] >> 16) | 0x0040).astype(np.uint16) + return out + + +def bf16_bits_to_float32(bits: ArrayLike) -> np.ndarray: + """bfloat16 bit patterns (uint16) -> float32 (exact).""" + b = np.ascontiguousarray(bits, dtype=np.uint16) + return (b.astype(np.uint32) << 16).view(np.float32) + + +def round_to_bf16(array: ArrayLike) -> np.ndarray: + """float32 values rounded to the nearest bfloat16 (RNE), returned as float32: the host emulation of a bf16 + upload, for building exact expectations in tests.""" + return bf16_bits_to_float32(float32_to_bf16_bits(array)) + + +# ----------------------------------------------------------------------------------------------- dtype maps + +_DTYPE_ALIASES = { + "bf16": "bfloat16", "bfloat16": "bfloat16", "fp32": "float32", "float32": "float32", "f32": "float32", + "bfp8": "bfloat8_b", "bfloat8_b": "bfloat8_b", "bf8": "bfloat8_b", "bfp4": "bfloat4_b", "bfloat4_b": "bfloat4_b", + "uint32": "uint32", "u32": "uint32", "int32": "int32", "i32": "int32", "uint16": "uint16", "u16": "uint16", + "uint8": "uint8", "u8": "uint8", +} + +# canonical numpy dtype used to build the host tensor for each ttnn dtype (bf16 / bfp go through float32) +_HOST_NUMPY = {"float32": np.float32, "bfloat16": np.float32, "bfloat8_b": np.float32, "bfloat4_b": np.float32, + "uint32": np.uint32, "int32": np.int32, "uint16": np.uint16, "uint8": np.uint8} + + +def ttnn_dtype(name_or_dtype: Any): + """``"bf16"`` / ``"fp32"`` / ``"bfp8"`` / ``"uint32"`` ... (or a ttnn dtype) -> ``ttnn.DataType``.""" + import ttnn + + if not isinstance(name_or_dtype, str): + return name_or_dtype + key = _DTYPE_ALIASES.get(name_or_dtype.strip().lower()) + if key is None: + raise ValueError(f"unknown dtype {name_or_dtype!r}; expected one of {sorted(_DTYPE_ALIASES)}") + return getattr(ttnn, key) + + +def dtype_name(dtype: Any) -> str: + """Canonical lower-case name of a ttnn dtype (``ttnn.bfloat16`` -> ``"bfloat16"``), or of an alias string.""" + if isinstance(dtype, str): + key = _DTYPE_ALIASES.get(dtype.strip().lower()) + if key is None: + raise ValueError(f"unknown dtype {dtype!r}") + return key + name = getattr(dtype, "name", None) or str(dtype).rsplit(".", 1)[-1] + name = name.lower() + return {"bfloat8_b": "bfloat8_b", "bfloat4_b": "bfloat4_b"}.get(name, name) + + +def _exact_torch(value: ArrayLike, dtype: Any): + """A torch tensor whose dtype matches ``dtype`` exactly for integers (ttnn would otherwise route int64 through + bf16) and is float32 / bf16 for float targets.""" + import torch + + name = dtype_name(dtype) + if isinstance(value, torch.Tensor): + value = value.detach().cpu() + if name == "bfloat16": + return value.to(torch.bfloat16) if value.dtype != torch.bfloat16 else value.contiguous() + if name in ("float32", "bfloat8_b", "bfloat4_b"): + return value.to(torch.float32).contiguous() + value = value.numpy() if value.dtype not in (torch.bfloat16,) else value.float().numpy() + a = np.asarray(value) + target = _HOST_NUMPY[name] + if np.issubdtype(np.dtype(target), np.integer): + if not (np.issubdtype(a.dtype, np.integer) or a.dtype == np.bool_): + raise TypeError(f"{name} tensor needs integer data, got {a.dtype}") + if a.size and np.issubdtype(a.dtype, np.signedinteger) and np.dtype(target).kind == "u" and a.min() < 0: + raise ValueError(f"negative values cannot be stored as {name}") + a = np.ascontiguousarray(a, dtype=target) + if not a.flags.writeable: # broadcast views are read-only; torch.from_numpy wants writable memory + a = a.copy() + t = torch.from_numpy(a) + return t.to(torch.bfloat16) if name == "bfloat16" else t + + +# ------------------------------------------------------------------------------------------- device side + +def to_host_tensor(value: ArrayLike, dtype: Any, layout: Any = None, *, shape: Optional[Sequence[int]] = None): + """numpy / torch / ttnn host tensor -> ttnn host tensor of ``dtype`` and ``layout`` (default ROW_MAJOR). + + A ttnn host tensor is checked (dtype, layout, optional shape) and returned unchanged. TILE layout tilizes on the + host: prefer ROW_MAJOR inputs plus an on-device ``ttnn.to_layout`` for large per-frame tensors.""" + import ttnn + + dtype = ttnn_dtype(dtype) + layout = ttnn.ROW_MAJOR_LAYOUT if layout is None else layout + if isinstance(value, ttnn.Tensor): + if value.storage_type() == ttnn.StorageType.DEVICE: + raise ValueError("expected a host tensor, got a device tensor") + if value.dtype != dtype or value.layout != layout: + raise ValueError(f"host tensor is {value.dtype}/{value.layout}, expected {dtype}/{layout}") + if shape is not None and tuple(value.shape) != tuple(shape): + raise ValueError(f"host tensor has shape {tuple(value.shape)}, expected {tuple(shape)}") + return value + t = _exact_torch(value, dtype) + if shape is not None and tuple(t.shape) != tuple(int(s) for s in shape): + raise ValueError(f"value has shape {tuple(t.shape)}, expected {tuple(shape)}") + return ttnn.from_torch(t, dtype=dtype, layout=layout) + + +def to_device(value: ArrayLike, device, dtype: Any = "bf16", layout: Any = None, memory_config: Any = None): + """Host data -> device tensor (``layout`` default TILE, ``memory_config`` default DRAM interleaved).""" + import ttnn + + layout = ttnn.TILE_LAYOUT if layout is None else layout + host = to_host_tensor(value, dtype, layout) + return ttnn.to_device(host, device, memory_config=memory_config or ttnn.DRAM_MEMORY_CONFIG) + + +def to_numpy(tensor: Any) -> np.ndarray: + """ttnn tensor (device or host) / torch tensor / array -> numpy. bf16 becomes float32 (exact); unsigned torch + dtypes without numpy support are converted through int64.""" + if isinstance(tensor, np.ndarray): + return tensor + try: + import ttnn + except ImportError: + ttnn = None + if ttnn is not None and isinstance(tensor, ttnn.Tensor): + tensor = ttnn.to_torch(tensor) + if hasattr(tensor, "detach") and hasattr(tensor, "cpu"): + import torch + + t = tensor.detach().cpu() + if t.dtype == torch.bfloat16: + return t.float().numpy() + try: + return t.numpy() + except TypeError: # torch.uint16 / uint32 on older numpy bridges + target = {torch.uint16: np.uint16, torch.uint32: np.uint32}.get(t.dtype, np.int64) + return t.to(torch.int64).numpy().astype(target) + return np.asarray(tensor) + + +def to_layout(tensor, layout): + """``ttnn.to_layout`` that returns the input unchanged when it already has ``layout`` (no extra program).""" + import ttnn + + return tensor if tensor.layout == layout else ttnn.to_layout(tensor, layout) + + +def typecast(tensor, dtype): + """``ttnn.typecast`` that is a no-op when ``tensor`` already has ``dtype``.""" + import ttnn + + dtype = ttnn_dtype(dtype) + return tensor if tensor.dtype == dtype else ttnn.typecast(tensor, dtype) + + +def to_fp32(tensor): + """Enter an fp32 island (typecast to FLOAT32 if needed).""" + return typecast(tensor, "float32") + + +def to_bf16(tensor): + """Leave an fp32 island (typecast to BFLOAT16 if needed).""" + return typecast(tensor, "bfloat16") + + +def fp32_island(fn: Callable[..., Any], *tensors, out_dtype: Any = "bfloat16"): + """Run ``fn`` on fp32 copies of ``tensors`` and cast its tensor output(s) to ``out_dtype`` (``None`` keeps + fp32). For numerically sensitive glue (box coding, softmax logits, geometry) between bf16 stages.""" + out = fn(*(to_fp32(t) for t in tensors)) + if out_dtype is None: + return out + if isinstance(out, (list, tuple)): + return type(out)(typecast(t, out_dtype) for t in out) + return typecast(out, out_dtype) + + +def upload_u8(array: ArrayLike, device, *, out_dtype: Any = "bfloat16", layout: Any = None, memory_config: Any = None): + """Upload uint8 data (4x less PCIe traffic than fp32) and typecast it on the device (exact for 0..255). + + In a traced model keep the uint8 tensor as the persistent trace input and make ``ttnn.typecast`` the first + traced op; this helper is the eager form.""" + import ttnn + + layout = ttnn.ROW_MAJOR_LAYOUT if layout is None else layout + u8 = ttnn.to_device(to_host_tensor(array, "uint8", layout), device, + memory_config=memory_config or ttnn.DRAM_MEMORY_CONFIG) + return typecast(u8, out_dtype) + + +class HostStaging: + """A persistent ROW_MAJOR host tensor written in place for every frame (RP section 2.11, rf-detr R3c). + + ``write(array)`` copies (and casts) into the tensor's own host buffer through a writable ``torch.from_dlpack`` + alias, so no ``ttnn.from_torch`` allocation happens per frame; ``.tensor`` is what you pass to + ``ttnn.copy_host_to_device_tensor`` (or ``TraceRunner.run(inputs=...)``). bf16 is written as RNE-rounded bit + patterns. If the alias check fails on some tt-metal build, ``zero_copy`` is ``False`` and ``write`` rebuilds the + host tensor with ``from_torch`` (same result, slower).""" + + _VIEW = {"float32": np.float32, "bfloat16": np.uint16, "uint32": np.uint32, "int32": np.int32, + "uint16": np.uint16, "uint8": np.uint8} + + def __init__(self, shape: Sequence[int], dtype: Any = "float32"): + import torch + import ttnn + + self.shape = tuple(int(s) for s in shape) + self.dtype = ttnn_dtype(dtype) + self._name = dtype_name(self.dtype) + if self._name not in self._VIEW: + raise ValueError(f"HostStaging supports {sorted(self._VIEW)}, not {self._name}") + zeros = np.zeros(self.shape, _HOST_NUMPY[self._name]) + self.tensor = ttnn.from_torch(_exact_torch(zeros, self.dtype), dtype=self.dtype, layout=ttnn.ROW_MAJOR_LAYOUT) + self.zero_copy = False + self._keep: tuple = () + try: + shard = self.tensor.host_buffer().get_shard(ttnn.MeshCoordinate(0, 0)) + raw = torch.from_dlpack(shard) + view = raw.numpy().view(self._VIEW[self._name]).reshape(self.shape) + probe = (np.arange(view.size) % 251 + 1).astype(view.dtype).reshape(self.shape) + view[...] = probe + back = to_numpy(self.tensor) + expect = bf16_bits_to_float32(probe) if self._name == "bfloat16" else probe + if np.array_equal(back.reshape(self.shape), expect): + self.zero_copy = True + self._keep = (shard, raw) + self.view = view + view[...] = 0 + except Exception: # noqa: BLE001 -- API drift: fall back to from_torch per write + self.zero_copy = False + if not self.zero_copy: + self.view = np.zeros(self.shape, self._VIEW[self._name]) + + def write(self, array: ArrayLike): + """Copy ``array`` (broadcastable to ``shape``) into the host buffer and return the host tensor.""" + a = np.asarray(array) + if self._name == "bfloat16": + self.view[...] = float32_to_bf16_bits(np.broadcast_to(a, self.shape)) + else: + self.view[...] = np.broadcast_to(a, self.shape).astype(self.view.dtype, copy=False) + if not self.zero_copy: + import ttnn + + data = bf16_bits_to_float32(self.view) if self._name == "bfloat16" else self.view + self.tensor = ttnn.from_torch(_exact_torch(data, self.dtype), dtype=self.dtype, + layout=ttnn.ROW_MAJOR_LAYOUT) + return self.tensor + + @property + def nbytes(self) -> int: + return int(math.prod(self.shape) * self.view.dtype.itemsize) diff --git a/code/tt_diffusion_planner/ttaw/trace.py b/code/tt_diffusion_planner/ttaw/trace.py new file mode 100644 index 0000000000000000000000000000000000000000..13f473e2297a2e232f5e83786d8e6df7b0d881ad --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/trace.py @@ -0,0 +1,1317 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C02: ``TraceRunner`` -- every device stage of a port runs inside metal traces. + +What it enforces (TT_PLATFORM.md section 3, REFERENCE_PATTERNS.md section 1.4): + +1. **Persistent I/O before any capture.** Inputs, RT-dev parameters, state buffers and (optionally) output buffers + are allocated (and given defined initial values: never uninitialised index buffers) when they are added, i.e. + before the first capture. A variant returns either the tensors its last ops produce (allocated during capture, + valid while the trace lives) or persistent outputs written with ``ctx.write_output`` (stable addresses shared by + several variants, never overwritten by another variant's replay). +2. **Warm-up, then capture.** Every variant runs eagerly ``warmup_runs`` times (kernel JIT, program cache, prepared + conv weights) before *any* trace is captured. Adding a variant after a capture releases all traces, warms the new + one and recaptures everything, because a warm-up after a capture could allocate into a trace's freed + intermediates. +3. **Strict capture.** ``device.set_program_cache_misses_allowed(False)`` during capture (a miss raises with the op + name instead of aborting with "Writes are not supported during trace capture"); ``end_trace_capture`` runs in a + ``finally`` block and a failed capture is released (an open capture left the process spinning in close_device: + RP section 1.4); the program-cache entry count must not change during capture. +4. **Variants.** Several traces keyed by name (shape buckets, modes, segments), chosen by the caller per run. +5. **RT-dev parameters.** Per-frame values that must not change the program (thresholds, timesteps, poses) live + in persistent device tensors refreshed before ``execute_trace``; unchanged values are not re-uploaded. +6. **1CQ / 2CQ.** With 2 CQs inputs are uploaded on CQ1 and ordered with events, exactly like + ``common/tools/check_dispatch.py`` (CQ1 waits for the last trace, uploads, records; CQ0 waits, replays, records). + ``stage_inputs=True`` uses the ``tt_cnn`` executor pattern instead: CQ1 writes a DRAM staging copy while the + previous trace still runs, and an eager copy on CQ0 moves it into the trace input before the replay. Every CQ0 + write into a trace input outside the protocol (:meth:`TraceRunner.write_input`) re-records the event CQ1 waits for. +7. **State inside the trace.** ``ctx.write_state(name, value)`` is ``ttnn.copy(value, buffer)`` into a persistent + buffer (FLOAT32 / UINT32 / BFLOAT16 / ... all supported), or no op at all when ``value`` was computed straight + into ``ctx.write_target(name)`` (``output_tensor=``). ``pingpong=True`` states use two buffers and two traces + per variant (phase 0 reads A writes B, phase 1 reads B writes A); the phase flips after every run of a variant + that writes the ping-pong states. Per-stream state (PLAN.md D16): ``add_state(..., banks=n)`` + + :meth:`TraceRunner.save_state` / :meth:`TraceRunner.load_state`, with the policy in :class:`StreamBanks`. +8. **Readback.** One packed output (:func:`pack_outputs`) gives one D2H; reads go into preallocated host tensors, + on CQ0 or (segmented pipelines, 2CQ) on CQ1 after a host-side event wait. +9. **Alloc tracking.** With ``TT_METAL_TRACE_ALLOC_TRACKING=1`` set before ``import ttnn``, ``ttnn.execute_trace`` + refuses to replay over live unsafe buffers; the runner acknowledges the outputs of traces captured after the + first one (they may be overwritten by an earlier trace's replay: read outputs before running another variant). + +Example:: + + runner = TraceRunner(device, num_command_queues=2) + runner.add_input("x", shape=(1, 1, 64, 64), dtype="bfloat16", layout=ttnn.TILE_LAYOUT) + runner.add_param("scale", 1.0) # fp32 [1,1,1,1], TILE + runner.add_variant("default", lambda ctx: ttnn.relu(ttnn.multiply(ctx["x"], ctx["scale"]))) + runner.capture() # warm-up, then capture + y = runner("default", inputs={"x": x_np}, params={"scale": 0.5}) # upload, replay, read -> numpy +""" +from __future__ import annotations + +import contextlib +import math +import time +from dataclasses import dataclass, field +from typing import Any, Callable, Dict, Iterator, List, Mapping, Optional, Sequence, Tuple + +import numpy as np + +from .io import InputError +from .tensors import TILE, dtype_name, round_up, to_host_tensor, to_numpy, ttnn_dtype + +__all__ = [ + "CQ_COMPUTE", + "CQ_INPUT", + "PackEntry", + "PackLayout", + "Packed", + "pack_outputs", + "SINGLE_ROW_MAX_ELEMS", + "PACK_ROW_ELEMS", + "TraceContext", + "TraceRunner", + "StreamBanks", + "alloc_tracking_enabled", +] + +CQ_COMPUTE = 0 # programs, traces and (by default) readback +CQ_INPUT = 1 # host -> device uploads when the device has 2 command queues + + +def alloc_tracking_enabled() -> bool: + """True when ``TT_METAL_TRACE_ALLOC_TRACKING=1`` was set before ttnn was imported (read from tt-metal).""" + try: + from ttnn.tools import trace_allocation_tracker as tracker + except ImportError: + return False + return bool(getattr(tracker, "TRACE_ALLOC_TRACKING", False)) + + +# ----------------------------------------------------------------------------------------------- packing + +# Packed layouts (PORT_LOG Q9 of the YOLOX port). A ROW_MAJOR tensor is stored page by page, one page per row, and +# the RM reshape / concat programs stage whole pages in L1 (reshape_rm_program_factory.cpp: 2 x the destination page +# per kernel copy when the pages are not 16-byte aligned), so a single-row pack of more than ~0.6 MB fails with +# "RM reshape dest staging does not fit in L1". Above SINGLE_ROW_MAX_ELEMS the pack is a [1, 1, rows, R] tensor of +# R = PACK_ROW_ELEMS elements per row: every RM page the packing programs touch stays below ROW_PAGE_MAX_BYTES, +# whatever the size of the outputs (tens of MB), and the readback is still ONE device-to-host copy. +SINGLE_ROW_MAX_ELEMS = 131072 # one [1, 1, 1, total] row up to 512 KiB of float32 (the 0.1.0 - 0.14.0 layout) +PACK_ROW_ELEMS = 1024 # elements per row of the multi-row layout (4 KiB float32 pages) +FLAT_MAX_BYTES = 64 << 10 # multi-row: tensors up to this size are flattened to one row, then cut into rows +ROW_PAGE_MAX_BYTES = 128 << 10 # multi-row: largest RM page (last dim x element size) a packed tensor may have +_ELEM_BYTES = {"float32": 4, "uint32": 4, "int32": 4, "bfloat16": 2, "uint16": 2, "uint8": 1} + + +@dataclass(frozen=True) +class PackEntry: + """One tensor inside a packed readback: elements ``[offset, offset + numel)`` reshaped to ``shape``. + + ``pitch > 0`` (multi-row layout only): the tensor's rows (its last dim, ``shape[-1]`` elements) + are stored ``pitch`` elements apart, zero-padded, i.e. elements ``[offset, offset + numel // shape[-1] * pitch)`` + viewed as ``(rows, pitch)`` hold it in their first ``shape[-1]`` columns. ``0``: contiguous.""" + + name: str + offset: int + numel: int + shape: Tuple[int, ...] + pitch: int = 0 + + def view(self, flat: np.ndarray) -> np.ndarray: + """This tensor inside the flat readback ``flat`` (a view, strided when ``pitch`` is set).""" + if not self.pitch: + return flat[self.offset:self.offset + self.numel].reshape(self.shape) + cols = int(self.shape[-1]) if self.shape else 1 + rows = self.numel // cols + return flat[self.offset:self.offset + rows * self.pitch].reshape(rows, self.pitch)[:, :cols].reshape( + self.shape) + + +@dataclass(frozen=True) +class PackLayout: + """Host-side description of a packed output (built at capture time, applied at every read). + + ``entries`` index the packed tensor's elements in row-major order, so :meth:`unpack` does not depend on the + device shape ``(1, 1, rows, row_elems)``: one row of ``total`` elements (``rows == 1``) or rows of + ``row_elems`` elements, every tensor starting on a row boundary (multi-row layout; an entry with a ``pitch`` + stores its rows zero-padded to that many elements, see :class:`PackEntry`).""" + + entries: Tuple[PackEntry, ...] + total: int + rows: int = 1 + row_elems: int = 0 # 0: one row of ``total`` elements + + @property + def shape(self) -> Tuple[int, int, int, int]: + """Shape of the packed device tensor.""" + return (1, 1, self.rows, self.row_elems or self.total) + + def unpack(self, flat: Any) -> Dict[str, np.ndarray]: + """Flat readback (any shape with ``total`` elements) -> ``{name: array}`` (views, no copies; the view of an + entry with a ``pitch`` is strided, not C-contiguous).""" + a = np.asarray(flat).reshape(-1) + if a.size != self.total: + raise ValueError(f"packed readback has {a.size} elements, layout expects {self.total}") + return {e.name: e.view(a) for e in self.entries} + + +@dataclass +class Packed: + """A packed device tensor plus its layout; return it from a variant function to get ``{name: array}`` back.""" + + tensor: Any + layout: PackLayout + + +def _elem_bytes(dtype: Any) -> int: + return _ELEM_BYTES.get(dtype_name(dtype), 4) + + +def _row_major(ttnn: Any, t: Any, dtype: Any) -> Any: + if t.dtype != dtype: + t = ttnn.typecast(t, dtype) + if t.layout != ttnn.ROW_MAJOR_LAYOUT: + t = ttnn.to_layout(t, ttnn.ROW_MAJOR_LAYOUT) + return t + + +def _flat_rows(ttnn: Any, t: Any, shape: Tuple[int, ...], numel: int, row_elems: int) -> Tuple[Any, int]: + """A small RM tensor -> one row -> zero-padded to whole rows -> ``[1, 1, k, row_elems]``.""" + padded = round_up(numel, row_elems) + if shape != (1, 1, 1, numel): + t = ttnn.reshape(t, (1, 1, 1, numel)) + if padded != numel: + t = ttnn.pad(t, [(0, 0), (0, 0), (0, 0), (0, padded - numel)], 0.0) + if padded != row_elems: + t = ttnn.reshape(t, (1, 1, padded // row_elems, row_elems)) + return t, padded + + +def _fill_rows(rows: int, cols: int, row_elems: int) -> int: + """Rows of ``cols`` elements a ``[rows, cols]`` tensor is zero-padded to so that it fills whole packed rows.""" + return round_up(rows, row_elems // math.gcd(cols, row_elems)) + + +def _row_pitches(cols: int, row_elems: int, elem: int) -> List[int]: + """Row pitches > ``cols`` worth trying: ``cols`` rounded up to every power of two dividing ``row_elems`` (rows of + that pitch fill a packed row every ``row_elems / gcd`` rows) and to ``row_elems`` (every row fills whole packed + rows), within the page budget.""" + steps, m = {row_elems}, 2 + while row_elems % m == 0: + steps.add(m) + m *= 2 + return sorted({p for p in (round_up(cols, m) for m in steps) if p != cols and p * elem <= ROW_PAGE_MAX_BYTES}) + + +def _pack_rows_segment(ttnn: Any, name: str, t: Any, shape: Tuple[int, ...], numel: int, row_elems: int, + elem: int) -> Tuple[Any, int, int]: + """One ROW_MAJOR tensor -> ``[1, 1, k, row_elems]`` holding its elements in order, zero-padded to whole rows. + Returns ``(segment, k * row_elems, pitch)`` (``pitch``: see :class:`PackEntry`). No RM page above + ``ROW_PAGE_MAX_BYTES`` is created or reshaped.""" + cols = int(shape[-1]) if shape else 1 + rows = numel // cols + if cols * elem <= ROW_PAGE_MAX_BYTES: + # [rows, cols] is a view of the RM tensor; rows * cols fills whole packed rows when rows % unit == 0 + rows_p = _fill_rows(rows, cols, row_elems) + if rows_p != rows and numel * elem <= FLAT_MAX_BYTES: + return (*_flat_rows(ttnn, t, shape, numel, row_elems), 0) # small: pad < row_elems elements, not rows + if rows_p != rows: + # Appending zero rows costs up to row_elems / gcd(cols, row_elems) - 1 rows: 80 MB for a [1, 1, 2, 20001] + # fp32 tensor. Take one flat row (when it fits the page budget) or rows zero-padded to a pitch instead + # when that is smaller by more than 1/8 of the tensor (typical shapes keep the contiguous layout). + options = [(p * _fill_rows(rows, p, row_elems), p) for p in _row_pitches(cols, row_elems, elem)] + if round_up(numel, row_elems) * elem <= ROW_PAGE_MAX_BYTES: + options.append((round_up(numel, row_elems), 0)) + if options and rows_p * cols - min(options)[0] > numel // 8: + size, pitch = min(options) # ties: the flat row, then the narrower pitch + if not pitch: + return (*_flat_rows(ttnn, t, shape, numel, row_elems), 0) + prows = size // pitch + if shape != (1, 1, rows, cols): + t = ttnn.reshape(t, (1, 1, rows, cols)) + t = ttnn.pad(t, [(0, 0), (0, 0), (0, prows - rows), (0, pitch - cols)], 0.0) + if pitch != row_elems: + t = ttnn.reshape(t, (1, 1, size // row_elems, row_elems)) + return t, size, pitch + # zero rows appended, then one RM reshape into rows of row_elems (source pages of cols elements, destination + # pages of row_elems elements: both small) + if shape != (1, 1, rows, cols): + t = ttnn.reshape(t, (1, 1, rows, cols)) + if rows_p != rows: + t = ttnn.pad(t, [(0, 0), (0, 0), (0, rows_p - rows), (0, 0)], 0.0) + if cols != row_elems: + t = ttnn.reshape(t, (1, 1, rows_p * cols // row_elems, row_elems)) + return t, rows_p * cols, 0 + if rows != 1: + raise ValueError(f"pack_outputs: {name!r} {shape} has rows of {cols * elem} B; the multi-row layout reads " + f"RM rows of at most {ROW_PAGE_MAX_BYTES} B: give it a narrower last dim (e.g. " + f"[1, 1, -1, {row_elems}]) before packing") + # one wide row (a flat vector): cut it into chunks of whole packed rows, each well inside the L1 budget + if shape != (1, 1, 1, numel): + t = ttnn.reshape(t, (1, 1, 1, numel)) + width = max(row_elems, (ROW_PAGE_MAX_BYTES // elem) // row_elems * row_elems) + parts, total = [], 0 + for start in range(0, numel, width): + stop = min(start + width, numel) + chunk = ttnn.slice(t, [0, 0, 0, start], [1, 1, 1, stop]) + chunk, padded = _flat_rows(ttnn, chunk, (1, 1, 1, stop - start), stop - start, row_elems) + parts.append(chunk) + total += padded + return (parts[0] if len(parts) == 1 else ttnn.concat(parts, dim=2)), total, 0 + + +def pack_outputs(tensors: Mapping[str, Any], *, dtype: Any = "float32", align: int = 32, + row_elems: Optional[int] = None) -> Packed: + """Pack several device tensors into ONE ROW_MAJOR tensor so the host reads them with one device-to-host copy. + + Each tensor is typecast to ``dtype`` (float32 default: exact for bf16 and for integers below 2**24) and converted + to ROW_MAJOR; call this inside the variant function (the ops become part of the trace) and return the result. + :meth:`PackLayout.unpack` (done by ``TraceRunner.read``) gives ``{name: array}`` back in the original shapes. + + Device layout: while the packed total is at most :data:`SINGLE_ROW_MAX_ELEMS`, one ``[1, 1, 1, total]`` row (each + tensor flattened and zero-padded to ``align`` elements, then concatenated): a ``[1, 1, N, 1]`` RM readback is N + pages and costs milliseconds, one row is one page (RP section 1.4). Above it (or with ``row_elems=R``) a + ``[1, 1, rows, R]`` tensor (default R = :data:`PACK_ROW_ELEMS`): each tensor is zero-padded to whole rows of R + elements and the segments are concatenated on the row dim, so no RM page exceeds :data:`ROW_PAGE_MAX_BYTES` + and outputs of tens of MB pack (a single row of more than ~0.6 MB does not fit the RM reshape's L1 staging: + YOLOX PORT_LOG Q9). A tensor is padded with zero rows (contiguous), or, when that wastes more than 1/8 of it + (an awkward last dim with few rows: e.g. ``[1, 1, 2, 20001]`` would take 1024 rows), flattened (up to the page + budget) or stored with its rows zero-padded to a wider pitch (:class:`PackEntry` ``pitch``). A last dim wider + than ``ROW_PAGE_MAX_BYTES`` is accepted for a flat vector (``[1, 1, 1, N]``, cut into chunks with + ``ttnn.slice``); any other tensor needs a narrower last dim (``ValueError``).""" + import ttnn + + if not tensors: + raise ValueError("pack_outputs needs at least one tensor") + if align <= 0: + raise ValueError("align must be positive") + if row_elems is not None and (int(row_elems) <= 0 or int(row_elems) % TILE): + raise ValueError(f"row_elems={row_elems}: expected a positive multiple of {TILE}") + dtype = ttnn_dtype(dtype) + elem = _elem_bytes(dtype) + items = [(str(name), t, tuple(int(s) for s in t.shape)) for name, t in tensors.items()] + empty = [name for name, _, shape in items if math.prod(shape) == 0] + if empty: + raise ValueError(f"pack_outputs: {empty} have no elements") + single_total = sum(round_up(math.prod(shape), align) for _, _, shape in items) + if row_elems is None and single_total <= SINGLE_ROW_MAX_ELEMS: + segments, entries, offset = [], [], 0 + for name, t, shape in items: # the single-row layout of ttaw 0.1.0 - 0.14.0 + numel = math.prod(shape) + t = _row_major(ttnn, t, dtype) + if shape != (1, 1, 1, numel): + t = ttnn.reshape(t, (1, 1, 1, numel)) + padded = round_up(numel, align) + if padded != numel: + t = ttnn.pad(t, [(0, 0), (0, 0), (0, 0), (0, padded - numel)], 0.0) + segments.append(t) + entries.append(PackEntry(name, offset, numel, shape)) + offset += padded + packed = segments[0] if len(segments) == 1 else ttnn.concat(segments, dim=-1) + return Packed(packed, PackLayout(tuple(entries), offset)) + r = int(row_elems or PACK_ROW_ELEMS) + segments, entries, offset = [], [], 0 + for name, t, shape in items: + numel = math.prod(shape) + t, padded, pitch = _pack_rows_segment(ttnn, name, _row_major(ttnn, t, dtype), shape, numel, r, elem) + segments.append(t) + entries.append(PackEntry(name, offset, numel, shape, pitch)) + offset += padded + packed = segments[0] if len(segments) == 1 else ttnn.concat(segments, dim=2) + return Packed(packed, PackLayout(tuple(entries), offset, rows=offset // r, row_elems=r)) + + +# --------------------------------------------------------------------------------------------- internals + +@dataclass +class _Slot: + """A persistent device tensor (input, parameter or state) allocated before any capture.""" + + name: str + kind: str # "input" | "param" | "state" | "output" + shape: Tuple[int, ...] + dtype: Any + layout: Any + memory_config: Any + buffers: List[Any] # 1 buffer, or 2 for a ping-pong state + init: Any # host tensor holding the initial value (states, params) / warm-up value (inputs) + staging: Any = None # DRAM staging copy (2CQ stage_inputs mode) + stage_fn: Optional[Callable[[Any, Any], Any]] = None + last_value: Optional[bytes] = None # params: bytes of the last uploaded value (skip unchanged uploads) + banks: List[Any] = field(default_factory=list) # states: save / load buffers (D16 per-stream state banks) + + @property + def pingpong(self) -> bool: + return len(self.buffers) == 2 + + def tensors(self) -> List[Any]: + """Every device tensor the slot owns.""" + return self.buffers + ([self.staging] if self.staging is not None else []) + self.banks + + +@dataclass +class _Variant: + name: str + fn: Callable[["TraceContext"], Any] + warmup_runs: int + + +@dataclass +class _Trace: + variant: str + phase: int + trace_id: Any + outputs: Any + leaves: List[Tuple[Tuple, Any, Optional[PackLayout]]] # (path, device tensor, pack layout or None) + steps: bool # writes the ping-pong states (flips the phase) + capture_ms: float + host_buffers: Optional[List[Any]] = None + done_event: Any = None + + +def _flatten(obj: Any, path: Tuple = ()) -> Iterator[Tuple[Tuple, Any]]: + """Leaves of a tensor / Packed / list / tuple / dict structure with their paths.""" + if isinstance(obj, Mapping): + for k, v in obj.items(): + yield from _flatten(v, path + (k,)) + elif isinstance(obj, (list, tuple)): + for i, v in enumerate(obj): + yield from _flatten(v, path + (i,)) + else: + yield path, obj + + +def _rebuild(obj: Any, values: Dict[Tuple, Any], path: Tuple = ()) -> Any: + if isinstance(obj, Mapping): + return {k: _rebuild(v, values, path + (k,)) for k, v in obj.items()} + if isinstance(obj, (list, tuple)): + return type(obj)(_rebuild(v, values, path + (i,)) for i, v in enumerate(obj)) + return values[path] + + +def _buffer_address(t: Any) -> Optional[int]: + try: + return int(t.buffer_address()) + except Exception: # noqa: BLE001 -- host tensor, deallocated tensor or a fake + return None + + +def _same_buffer(a: Any, b: Any) -> bool: + """``a`` is ``b`` or a tensor over the same device buffer.""" + if a is b: + return True + addr = _buffer_address(a) + return addr is not None and addr == _buffer_address(b) + + +def _check_copy(what: str, value: Any, slot: "_Slot") -> None: + """The preconditions of ``ttnn.copy(value, )``, checked in Python so a mistake is a clear error + instead of a TT_FATAL inside an open capture: same logical shape and layout; a dtype change only in TILE.""" + import ttnn + + shape = tuple(int(s) for s in value.shape) + if shape != slot.shape: + raise ValueError(f"{what}: value shape {shape} != {slot.shape}") + if value.layout != slot.layout: + raise ValueError(f"{what}: value layout {value.layout} != {slot.layout} (ttnn.copy keeps the layout)") + if value.dtype != slot.dtype and slot.layout != ttnn.TILE_LAYOUT: + raise ValueError(f"{what}: dtype {value.dtype} -> {slot.dtype} needs TILE layout (ttnn.copy)") + + +class TraceContext(Mapping): + """What a variant function receives: persistent tensors by name (``ctx["x"]``), and state writes. + + ``ctx[name]`` returns an input, a parameter, a persistent output buffer, or the buffer a state is *read* from in + this phase. ``ctx.write_state(name, value)`` copies ``value`` into the buffer the state is *written* to (in-place + states: the same buffer, so read everything you need from it before writing); ``ctx.write_output(name, value)`` + copies into a persistent output and returns it.""" + + def __init__(self, runner: "TraceRunner", variant: str, phase: int, capturing: bool): + self._runner = runner + self.variant = variant + self.phase = phase + self.capturing = capturing + self.writes: set = set() + + @property + def device(self): + return self._runner.device + + def __getitem__(self, name: str): + slot = self._runner._slot(name) + if slot.kind == "state": + return self.state(name) + return slot.buffers[0] + + def __iter__(self) -> Iterator[str]: + return iter(self._runner._slots) + + def __len__(self) -> int: + return len(self._runner._slots) + + def state(self, name: str): + """The buffer state ``name`` is read from in this phase.""" + slot = self._runner._slot(name, "state") + return slot.buffers[self.phase] if slot.pingpong else slot.buffers[0] + + def write_target(self, name: str): + """The buffer state ``name`` is written to in this phase (ping-pong: the other buffer; in place: the same + one). Pass it as ``output_tensor=`` of the op that produces the new state and then call + ``write_state(name, it)``: no copy program is traced (probe P13: saves one program per state per frame).""" + slot = self._runner._slot(name, "state") + return slot.buffers[1 - self.phase] if slot.pingpong else slot.buffers[0] + + def write_state(self, name: str, value) -> None: + """``ttnn.copy(value, )``: same logical shape and layout as the state (dtype may differ in + TILE layout). The copy is part of the trace; it is skipped when ``value`` already is the write buffer + (:meth:`write_target`).""" + import ttnn + + slot = self._runner._slot(name, "state") + target = self.write_target(name) + _check_copy(f"state {name!r}", value, slot) + if not _same_buffer(value, target): + ttnn.copy(value, target) + self.writes.add(name) + + def write_output(self, name: str, value): + """``ttnn.copy(value, )`` (part of the trace); returns the output buffer, which the + variant can return as (part of) its outputs. Skipped when ``value`` already is that buffer.""" + import ttnn + + slot = self._runner._slot(name, "output") + _check_copy(f"output {name!r}", value, slot) + if not _same_buffer(value, slot.buffers[0]): + ttnn.copy(value, slot.buffers[0]) + return slot.buffers[0] + + +class TraceRunner: + """Persistent device I/O + warm-up + capture + replay of one model's traced stages (see the module docstring). + + Args: + device: an open ttnn device. + num_command_queues: 1 or 2. ``None`` uses what :func:`.device.open_device` recorded (1 if unknown). + warmup_runs: eager runs of each variant (and phase) before any capture. + stage_inputs: 2CQ only: upload into DRAM staging buffers and copy them into the trace inputs with an eager + op on CQ0 (the upload of frame k+1 then overlaps the replay of frame k). + forbid_cache_misses: call ``device.set_program_cache_misses_allowed(False)`` during capture. + alloc_tracking: ``True`` raises unless the process tracks trace allocations + (``TT_METAL_TRACE_ALLOC_TRACKING=1`` set before ``import ttnn``). Whenever tracking is active, the outputs + of traces captured after the first are acknowledged as corruptible (see the module docstring). + name: label used in messages and ``describe()``. + """ + + def __init__(self, device, *, num_command_queues: Optional[int] = None, warmup_runs: int = 1, + stage_inputs: bool = False, forbid_cache_misses: bool = True, alloc_tracking: Optional[bool] = None, + name: str = "model"): + from .device import open_info + + opened = open_info(device).get("num_command_queues") + if num_command_queues is None: + num_command_queues = opened or 1 + if num_command_queues not in (1, 2): + raise ValueError(f"num_command_queues={num_command_queues}: expected 1 or 2") + if opened is not None and num_command_queues > opened: + raise ValueError(f"the device was opened with {opened} command queue(s); cannot run {num_command_queues}") + if stage_inputs and num_command_queues != 2: + raise ValueError("stage_inputs=True needs num_command_queues=2") + if warmup_runs < 1: + raise ValueError("warmup_runs must be >= 1 (capture needs a warm program cache)") + tracking = alloc_tracking_enabled() + if alloc_tracking and not tracking: + raise RuntimeError("alloc_tracking=True but trace allocation tracking is off: export " + "TT_METAL_TRACE_ALLOC_TRACKING=1 (optionally TT_METAL_TRACE_ALLOC_TRACEBACKS=1) " + "before Python imports ttnn") + self.device = device + self.name = name + self.num_command_queues = int(num_command_queues) + self.warmup_runs = int(warmup_runs) + self.stage_inputs = bool(stage_inputs) + self.forbid_cache_misses = bool(forbid_cache_misses) + self.alloc_tracking = bool(tracking) + self._slots: Dict[str, _Slot] = {} + self._variants: Dict[str, _Variant] = {} + self._traces: Dict[Tuple[str, int], _Trace] = {} + self._phase = 0 + self._last_phase: Dict[str, int] = {} + self._last_variant: Optional[str] = None + self._pending_params: Dict[str, Any] = {} + self._op_event = None + self._stage_free_event = None + self._captures = 0 + self._capturing = False + self._eager_warmups: List[Callable[[], Any]] = [] + self._eager_warmed = False + self._closed = False + self.timings_ms: Dict[str, Dict[str, float]] = {} + if hasattr(device, "enable_program_cache"): + device.enable_program_cache() + + # ------------------------------------------------------------------------------------- registration + def _check_registration(self, what: str) -> None: + self._check_open() + if self._capturing: + raise RuntimeError(f"cannot add {what} while a variant is being warmed up or captured: register " + "inputs, params, states, outputs and variants before capture()") + + def _new_slot(self, name: str, kind: str, init: Any, shape: Optional[Sequence[int]], dtype: Any, layout: Any, + memory_config: Any, n_buffers: int, stage_fn: Optional[Callable], n_banks: int = 0) -> _Slot: + import ttnn + + self._check_registration(f"{kind} {name!r}") + if name in self._slots: + raise ValueError(f"{name!r} is already registered (as {self._slots[name].kind})") + if self._traces: + raise RuntimeError(f"cannot add {kind} {name!r} after capture: persistent tensors must exist before the " + "first capture (release() and rebuild)") + layout = ttnn.ROW_MAJOR_LAYOUT if layout is None else layout + memory_config = ttnn.DRAM_MEMORY_CONFIG if memory_config is None else memory_config + dtype = ttnn_dtype(dtype) + if init is None: + if shape is None: + raise ValueError(f"{kind} {name!r}: give init= or shape=") + init = np.zeros(tuple(shape), np.float32 if dtype_name(dtype) in ("float32", "bfloat16", "bfloat8_b", + "bfloat4_b") else np.int64) + if isinstance(init, ttnn.Tensor): + host = to_host_tensor(init, dtype, layout, shape=shape) + else: + arr = to_numpy(init) if hasattr(init, "detach") else np.asarray(init) + if shape is not None: + arr = np.broadcast_to(arr, tuple(shape)) + host = to_host_tensor(arr, dtype, layout) + shape_t = tuple(int(s) for s in host.shape) + + def allocate(config: Any): + buf = ttnn.allocate_tensor_on_device(ttnn.Shape(list(shape_t)), dtype, layout, self.device, config) + ttnn.copy_host_to_device_tensor(host, buf, cq_id=CQ_COMPUTE) # defined contents, never garbage + return buf + + buffers = [allocate(memory_config) for _ in range(n_buffers)] + staging = allocate(ttnn.DRAM_MEMORY_CONFIG) if self.stage_inputs and kind in ("input", "param") else None + banks = [allocate(ttnn.DRAM_MEMORY_CONFIG) for _ in range(n_banks)] + slot = _Slot(name, kind, shape_t, dtype, layout, memory_config, buffers, host, staging, + stage_fn or (lambda src, dst: ttnn.copy(src, dst)), banks=banks) + self._slots[name] = slot + return slot + + def add_input(self, name: str, init: Any = None, *, shape: Optional[Sequence[int]] = None, + dtype: Any = "bfloat16", layout: Any = None, memory_config: Any = None, + stage_fn: Optional[Callable[[Any, Any], Any]] = None): + """A persistent trace input (default ROW_MAJOR in DRAM). ``init`` (array / torch / ttnn host tensor) is the + initial content and the warm-up input; zeros of ``shape`` otherwise -- give real data when the graph + gathers with these values (garbage indices can hang the chip). ``stage_fn(staging, persistent)`` replaces + the eager ``ttnn.copy`` in ``stage_inputs`` mode (e.g. a reshard into a sharded L1 input). Returns the + device tensor.""" + return self._new_slot(name, "input", init, shape, dtype, layout, memory_config, 1, stage_fn).buffers[0] + + def add_param(self, name: str, value: Any = 0.0, *, shape: Sequence[int] = (1, 1, 1, 1), dtype: Any = "float32", + layout: Any = None, memory_config: Any = None): + """An RT-dev parameter: a persistent device tensor (default fp32 ``[1, 1, 1, 1]`` TILE, which broadcasts in + ttnn binary ops) refreshed before a replay when ``run(params=...)`` / :meth:`set_params` changes it.""" + import ttnn + + layout = ttnn.TILE_LAYOUT if layout is None else layout + slot = self._new_slot(name, "param", np.broadcast_to(np.asarray(value), tuple(shape)), tuple(shape), dtype, + layout, memory_config, 1, None) + slot.last_value = self._param_bytes(slot, value) + return slot.buffers[0] + + def add_state(self, name: str, init: Any = None, *, shape: Optional[Sequence[int]] = None, dtype: Any = "float32", + layout: Any = None, memory_config: Any = None, pingpong: bool = False, banks: int = 0): + """Temporal state kept on the device across replays (memory queues, previous BEV, ring buffers). + + In-place (default): one buffer, read via ``ctx[name]`` and written via ``ctx.write_state``. ``pingpong``: + two buffers and two traces per variant. ``init`` is restored by :meth:`reset_state`. ``banks``: extra DRAM + buffers of the same spec for :meth:`save_state` / :meth:`load_state` (one per stream id of a + :class:`StreamBanks`, PLAN.md D16), allocated now because nothing may be allocated after a capture. + Returns the buffer(s).""" + import ttnn + + if banks < 0: + raise ValueError("banks must be >= 0") + layout = ttnn.TILE_LAYOUT if layout is None else layout + slot = self._new_slot(name, "state", init, shape, dtype, layout, memory_config, 2 if pingpong else 1, None, + n_banks=int(banks)) + return tuple(slot.buffers) if pingpong else slot.buffers[0] + + def add_output(self, name: str, *, shape: Sequence[int], dtype: Any = "float32", layout: Any = None, + memory_config: Any = None): + """A persistent output buffer (default TILE in DRAM, zeros) allocated before any capture. Variants write it + with ``ctx.write_output(name, value)`` (one traced ``ttnn.copy``) and return it; its address survives + recaptures and is shared by every variant (e.g. shape buckets with one readback). Returns the buffer.""" + import ttnn + + layout = ttnn.TILE_LAYOUT if layout is None else layout + return self._new_slot(name, "output", None, shape, dtype, layout, memory_config, 1, None).buffers[0] + + def add_variant(self, name: str, fn: Callable[[TraceContext], Any], *, warmup_runs: Optional[int] = None) -> None: + """Register a traced function. ``fn(ctx)`` runs ttnn ops on ``ctx[...]`` tensors and returns a device tensor, + a :class:`Packed`, or a list / tuple / dict of them (the persistent trace outputs). No host I/O, no + ``synchronize``, no torch ops inside ``fn``: it is called for warm-up and then recorded.""" + self._check_registration(f"variant {name!r}") + if name in self._variants: + raise ValueError(f"variant {name!r} already exists") + self._variants[name] = _Variant(name, fn, self.warmup_runs if warmup_runs is None else int(warmup_runs)) + + def add_eager_warmup(self, fn: Callable[[], Any]) -> None: + """Register eager device work that the model runs *between* replays (host-fallback glue, an eager layout + change of a read-back tensor, ...): ``fn()`` runs once in the first :meth:`capture`, before any trace is + captured, so its programs are compiled -- and their kernel binaries allocated in DRAM -- before the first + capture. A program compiled after a capture shares the address space of the traces' freed intermediates and + a replay can overwrite its binaries (tt-metal ``tech_reports/.../TraceCorrectness.md``: corruption or a + hang); ``TT_METAL_TRACE_ALLOC_TRACKING=1`` reports it.""" + self._check_registration("an eager warm-up") + if self._traces or self._eager_warmed: + raise RuntimeError("add eager warm-ups before the first capture (release() and rebuild)") + self._eager_warmups.append(fn) + + # ------------------------------------------------------------------------------------------ capture + @property + def phases(self) -> int: + """2 when a ping-pong state exists (two traces per variant), else 1.""" + return 2 if any(s.pingpong for s in self._slots.values()) else 1 + + @property + def captured(self) -> bool: + return bool(self._variants) and all((v, p) in self._traces for v in self._variants for p in range(self.phases)) + + def capture(self) -> None: + """Warm up and capture every registered variant that has no trace yet (idempotent). If traces already exist + and a new variant was added, all traces are released, the new variants warmed up, and all recaptured.""" + import ttnn + + self._check_open() + if not self._variants: + raise RuntimeError("no variant registered (add_variant)") + pending = [v for v in self._variants if any((v, p) not in self._traces for p in range(self.phases))] + if not pending: + return + if self._traces: + ttnn.synchronize_device(self.device) # no replay of a trace being released is still in flight + self._release_traces() + warm, pending = pending, list(self._variants) + else: + warm = pending + self._capturing = True + try: + if not self._eager_warmed: + self._warm_eager_programs() + for vname in warm: + t0 = time.perf_counter() + variant = self._variants[vname] + for phase in range(self.phases): + for _ in range(variant.warmup_runs): + ctx = TraceContext(self, vname, phase, capturing=False) + outputs = variant.fn(ctx) + ttnn.synchronize_device(self.device) + self._free_transient(outputs) + self.timings_ms.setdefault(vname, {})["warmup"] = (time.perf_counter() - t0) * 1e3 + self._phase = 0 + self.reset_state() + for vname in pending: + self.timings_ms.setdefault(vname, {})["capture"] = 0.0 + for phase in range(self.phases): + self._capture_one(vname, phase) + finally: + self._capturing = False + ttnn.synchronize_device(self.device) + if self.num_command_queues == 2: + self._op_event = ttnn.record_event(self.device, CQ_COMPUTE) + self._stage_free_event = self._op_event + + def _warm_eager_programs(self) -> None: + """Compile the runner's own eager programs before the first capture: the staging copies (``stage_inputs``), + the state <-> bank copies (``banks``) and the registered eager warm-ups. Contents are kept: a staging buffer + mirrors its input (:meth:`write_input` writes both), and a state and its first bank hold the same value + before the states are reset ahead of the captures.""" + import ttnn + + for slot in self._slots.values(): + if slot.staging is not None: + slot.stage_fn(slot.staging, slot.buffers[0]) + if slot.banks: + ttnn.copy(slot.buffers[0], slot.banks[0]) + ttnn.copy(slot.banks[0], slot.buffers[0]) + for fn in self._eager_warmups: + fn() + ttnn.synchronize_device(self.device) + self._eager_warmed = True + + @contextlib.contextmanager + def _eager(self, what: str) -> Iterator[None]: + """Eager runner work after a capture must not compile a program (see :meth:`add_eager_warmup`).""" + dev = self.device + count = getattr(dev, "num_program_cache_entries", None) + before = count() if (self._traces and count is not None) else None + yield + if before is not None and count() != before: + raise RuntimeError(f"{self.name}: {what} compiled a new program after capture; its kernel binaries share " + "DRAM with the traces' freed intermediates, so a replay can overwrite them. Run it " + "before the first capture (add_eager_warmup) or keep the tensor specs of the warmed " + "program") + + def _capture_one(self, vname: str, phase: int) -> None: + import ttnn + + dev = self.device + variant = self._variants[vname] + ctx = TraceContext(self, vname, phase, capturing=True) + entries_before = dev.num_program_cache_entries() if hasattr(dev, "num_program_cache_entries") else None + forbid = self.forbid_cache_misses and hasattr(dev, "set_program_cache_misses_allowed") + t0 = time.perf_counter() + if forbid: + dev.set_program_cache_misses_allowed(False) + trace_id = ttnn.begin_trace_capture(dev, cq_id=CQ_COMPUTE) + ok = ended = False + try: + outputs = variant.fn(ctx) + ok = True + finally: + try: + ttnn.end_trace_capture(dev, trace_id, cq_id=CQ_COMPUTE) + ended = True + finally: + if forbid: + dev.set_program_cache_misses_allowed(True) + if not (ok and ended): + self._safe_release(trace_id) + try: + entries_after = dev.num_program_cache_entries() if entries_before is not None else None + if entries_before is not None and entries_after != entries_before: + raise RuntimeError(f"{self.name}/{vname}: the program cache grew during capture ({entries_before} -> " + f"{entries_after}); warm-up does not cover the traced graph") + leaves = [] + for path, leaf in _flatten(outputs): + if isinstance(leaf, Packed): + leaves.append((path, leaf.tensor, leaf.layout)) + elif isinstance(leaf, ttnn.Tensor): + leaves.append((path, leaf, None)) + else: + raise TypeError(f"{self.name}/{vname}: output {path} is {type(leaf).__name__}, expected a device " + "tensor or Packed") + if not leaves: + raise ValueError(f"{self.name}/{vname}: the variant returned no output tensor") + pingpong = {s.name for s in self._slots.values() if s.pingpong} + written = ctx.writes & pingpong + if written and written != pingpong: + raise RuntimeError(f"{self.name}/{vname}: writes ping-pong states {sorted(written)} but not " + f"{sorted(pingpong - written)}; a stepping variant must write all of them") + except BaseException: + self._safe_release(trace_id) + raise + if self.alloc_tracking and self._captures > 0: + from ttnn.tools import trace_allocation_tracker as tracker + + keep = self._persistent_addresses() + for _, tensor, _ in leaves: + if _buffer_address(tensor) not in keep: + tracker.acknowledge_corruptible(tensor) + self._captures += 1 + ms = (time.perf_counter() - t0) * 1e3 + self._traces[(vname, phase)] = _Trace(vname, phase, trace_id, outputs, leaves, bool(written), ms) + self.timings_ms[vname]["capture"] += ms + + def _persistent_addresses(self) -> set: + addrs = set() + for slot in self._slots.values(): + for t in slot.tensors(): + a = _buffer_address(t) + if a is not None: + addrs.add(a) + return addrs + + def _deallocate_outputs(self, tensors: Iterator[Any]) -> None: + """Deallocate op-produced output tensors, never a persistent buffer or a view of one. ``force=False``: a + tensor sharing its device memory with another owner (a view of a model weight returned as an output) is + left to its owners -- ``ttnn.deallocate`` forces by default and would free the weight.""" + import ttnn + + keep = self._persistent_addresses() + seen = set() + for tensor in tensors: + if not isinstance(tensor, ttnn.Tensor) or id(tensor) in seen: + continue + seen.add(id(tensor)) + if tensor.is_allocated() and _buffer_address(tensor) not in keep: + ttnn.deallocate(tensor, False) + + def _free_transient(self, outputs: Any) -> None: + """Deallocate warm-up / eager outputs (see :meth:`_deallocate_outputs`).""" + self._deallocate_outputs(leaf.tensor if isinstance(leaf, Packed) else leaf for _, leaf in _flatten(outputs)) + + def _safe_release(self, trace_id) -> None: + import ttnn + + try: + ttnn.release_trace(self.device, trace_id) + except Exception: # noqa: BLE001 -- best effort on an error path; the original error is re-raised + pass + + # -------------------------------------------------------------------------------------------- inputs + def _slot(self, name: str, kind: Optional[str] = None) -> _Slot: + slot = self._slots.get(name) + if slot is None: + raise KeyError(f"{self.name}: no input / param / state named {name!r}") + if kind is not None and slot.kind != kind: + raise KeyError(f"{self.name}: {name!r} is a {slot.kind}, not a {kind}") + return slot + + @staticmethod + def _param_array(slot: _Slot, value: Any) -> np.ndarray: + dt = np.float32 if dtype_name(slot.dtype) in ("float32", "bfloat16", "bfloat8_b", "bfloat4_b") else np.int64 + return np.ascontiguousarray(np.broadcast_to(np.asarray(value, dtype=dt), slot.shape)) + + def _param_bytes(self, slot: _Slot, value: Any) -> bytes: + return self._param_array(slot, value).tobytes() + + def set_params(self, **values: Any) -> None: + """Queue RT-dev parameter values for the next run (uploaded only if they changed).""" + for name in values: + self._slot(name, "param") + self._pending_params.update(values) + + def _collect_uploads(self, inputs: Optional[Mapping[str, Any]], + params: Optional[Mapping[str, Any]]) -> List[Tuple[_Slot, Any, Optional[bytes]]]: + uploads = [] + for name, value in (inputs or {}).items(): + slot = self._slot(name, "input") + uploads.append((slot, to_host_tensor(value, slot.dtype, slot.layout, shape=slot.shape), None)) + merged = dict(self._pending_params) + merged.update(params or {}) + for name, value in merged.items(): + slot = self._slot(name, "param") + arr = self._param_array(slot, value) + key = arr.tobytes() + if key != slot.last_value: + uploads.append((slot, to_host_tensor(arr, slot.dtype, slot.layout), key)) + return uploads + + def _ensure_events(self) -> None: + """2CQ: the CQ0 events CQ1 waits for exist (``capture()`` records them; a partially failed capture, which + leaves earlier variants runnable, does not get that far).""" + import ttnn + + if self._op_event is None: + self._op_event = ttnn.record_event(self.device, CQ_COMPUTE) + if self._stage_free_event is None: + self._stage_free_event = self._op_event + + def _enqueue_uploads(self, uploads: List[Tuple[_Slot, Any, Optional[bytes]]]) -> None: + import ttnn + + if uploads: + if self.num_command_queues == 1: + for slot, host, _ in uploads: + ttnn.copy_host_to_device_tensor(host, slot.buffers[0], cq_id=CQ_COMPUTE) + elif self.stage_inputs: + self._ensure_events() + ttnn.wait_for_event(CQ_INPUT, self._stage_free_event) + for slot, host, _ in uploads: + ttnn.copy_host_to_device_tensor(host, slot.staging, cq_id=CQ_INPUT) + written = ttnn.record_event(self.device, CQ_INPUT) + ttnn.wait_for_event(CQ_COMPUTE, written) + with self._eager("a stage_fn copy"): + for slot, _, _ in uploads: + slot.stage_fn(slot.staging, slot.buffers[0]) + self._stage_free_event = ttnn.record_event(self.device, CQ_COMPUTE) + else: + self._ensure_events() + ttnn.wait_for_event(CQ_INPUT, self._op_event) + for slot, host, _ in uploads: + ttnn.copy_host_to_device_tensor(host, slot.buffers[0], cq_id=CQ_INPUT) + written = ttnn.record_event(self.device, CQ_INPUT) + ttnn.wait_for_event(CQ_COMPUTE, written) + for slot, _, key in uploads: + if key is not None: + slot.last_value = key + self._pending_params.clear() + + def write_input(self, name: str, value: Any) -> None: + """Upload ``value`` into input ``name`` now (CQ0, outside any trace): e.g. a realistic warm-up sample. + With 2 CQs the event that later CQ1 uploads wait for is re-recorded after this write, so an upload of the + next frame can never land before it (both write the same buffer from different queues).""" + import ttnn + + self._check_open() + slot = self._slot(name, "input") + host = to_host_tensor(value, slot.dtype, slot.layout, shape=slot.shape) + ttnn.copy_host_to_device_tensor(host, slot.buffers[0], cq_id=CQ_COMPUTE) + if slot.staging is not None: # the staging buffer mirrors the input (the stage-copy warm-up keeps it) + ttnn.copy_host_to_device_tensor(host, slot.staging, cq_id=CQ_COMPUTE) + if self.num_command_queues == 2 and self._op_event is not None: + self._op_event = ttnn.record_event(self.device, CQ_COMPUTE) + if self.stage_inputs: + self._stage_free_event = self._op_event + + # --------------------------------------------------------------------------------------------- run + def _trace_for(self, variant: Optional[str], phase: Optional[int] = None) -> _Trace: + self._check_open() + name = variant or self._last_variant or (next(iter(self._variants)) if len(self._variants) == 1 else None) + if name is None: + raise ValueError("variant name required") + if name not in self._variants: + raise KeyError(f"{self.name}: unknown variant {name!r}; have {sorted(self._variants)}") + p = self._phase if phase is None else phase + trace = self._traces.get((name, p)) + if trace is None: + raise RuntimeError(f"{self.name}: variant {name!r} is not captured; call capture() first") + return trace + + def _execute(self, trace: _Trace) -> None: + import ttnn + + ttnn.execute_trace(self.device, trace.trace_id, cq_id=CQ_COMPUTE, blocking=False) + if self.num_command_queues == 2: + self._op_event = ttnn.record_event(self.device, CQ_COMPUTE) + trace.done_event = self._op_event + self._last_variant = trace.variant + self._last_phase[trace.variant] = trace.phase + if trace.steps: + self._phase ^= 1 + + def upload(self, inputs: Optional[Mapping[str, Any]] = None, params: Optional[Mapping[str, Any]] = None) -> int: + """Enqueue the uploads of ``inputs`` and changed ``params`` (CQ0, or CQ1 + events with 2 CQs) without + replaying anything; the next :meth:`run` / :meth:`replay` consumes them. Returns the number of tensors + uploaded.""" + self._check_open() + if not self.captured: + raise RuntimeError(f"{self.name}: capture() before uploading") + uploads = self._collect_uploads(inputs, params) + self._enqueue_uploads(uploads) + return len(uploads) + + def run(self, variant: Optional[str] = None, inputs: Optional[Mapping[str, Any]] = None, + params: Optional[Mapping[str, Any]] = None): + """Upload ``inputs`` (name -> numpy / torch / ttnn host tensor) and changed ``params``, then replay the + variant's trace without blocking. Returns its device outputs (valid once the replay finished).""" + trace = self._trace_for(variant) + self._enqueue_uploads(self._collect_uploads(inputs, params)) + self._execute(trace) + return trace.outputs + + def replay(self, variant: Optional[str] = None, n: int = 1) -> None: + """Replay ``n`` times with no uploads (back-to-back device timing); honours ping-pong phases.""" + for _ in range(n): + self._execute(self._trace_for(variant)) + + def outputs(self, variant: Optional[str] = None, phase: Optional[int] = None): + """Device outputs of a variant's trace (default: the phase that ran last for it, else phase 0).""" + name = variant or self._last_variant + if phase is None and name is not None: + phase = self._last_phase.get(name, 0) + return self._trace_for(name, phase if phase is not None else 0).outputs + + def read(self, variant: Optional[str] = None, *, cq_id: int = CQ_COMPUTE, as_torch: bool = False): + """Read the outputs of the last run of ``variant`` into preallocated host tensors (blocking). + + Returns the output structure with numpy arrays (``as_torch=True``: torch tensors); a :class:`Packed` leaf + becomes ``{name: array}``. ``cq_id=1`` (2CQ) waits on the host for the trace's completion event and reads on + CQ1, so CQ0 can already replay the next segment (segmented D2H).""" + import ttnn + + name = variant or self._last_variant + if name is None or name not in self._last_phase: + raise RuntimeError(f"{self.name}: variant {name!r} has not run yet") + trace = self._trace_for(name, self._last_phase[name]) + if cq_id not in (CQ_COMPUTE, CQ_INPUT): + raise ValueError(f"cq_id={cq_id}: expected 0 or 1") + if cq_id == CQ_INPUT: + if self.num_command_queues != 2: + raise ValueError("cq_id=1 needs a device opened with 2 command queues") + ttnn.event_synchronize(trace.done_event) + if trace.host_buffers is None: + trace.host_buffers = [ttnn.allocate_tensor_on_host(t.spec, self.device) for _, t, _ in trace.leaves] + values: Dict[Tuple, Any] = {} + for (path, tensor, layout), host in zip(trace.leaves, trace.host_buffers): + ttnn.copy_device_to_host_tensor(tensor, host, blocking=True, cq_id=cq_id) + arr = to_numpy(host) + value: Any = layout.unpack(arr) if layout is not None else arr + if as_torch: + import torch + + value = ({k: torch.from_numpy(np.ascontiguousarray(v)) for k, v in value.items()} + if isinstance(value, dict) else torch.from_numpy(np.ascontiguousarray(value))) + values[path] = value + structure = trace.outputs + if isinstance(structure, Packed) or not isinstance(structure, (Mapping, list, tuple)): + return values[()] + return _rebuild(structure, values) + + def __call__(self, variant: Optional[str] = None, inputs: Optional[Mapping[str, Any]] = None, + params: Optional[Mapping[str, Any]] = None, *, as_torch: bool = False): + """``run`` + ``read`` (the common synchronous path).""" + self.run(variant, inputs, params) + return self.read(variant, as_torch=as_torch) + + def run_eager(self, variant: Optional[str] = None, inputs: Optional[Mapping[str, Any]] = None, + params: Optional[Mapping[str, Any]] = None): + """Upload, run the variant's function *eagerly* (no trace) and return its outputs read to numpy, freeing + every eager buffer before returning (safe next to captured traces). State writes happen as in a replay. + Use it for replay-vs-eager bit checks.""" + import ttnn + + self._check_open() + name = variant or self._last_variant or (next(iter(self._variants)) if self._variants else None) + if name not in self._variants: + raise KeyError(f"{self.name}: unknown variant {name!r}; have {sorted(self._variants)}") + variant_obj = self._variants[name] + uploads = self._collect_uploads(inputs, params) + for slot, host, key in uploads: + ttnn.copy_host_to_device_tensor(host, slot.buffers[0], cq_id=CQ_COMPUTE) + if key is not None: + slot.last_value = key + self._pending_params.clear() + ctx = TraceContext(self, name, self._phase, capturing=False) + with self._eager(f"run_eager({name!r})"): + outputs = variant_obj.fn(ctx) + values = {} + for path, leaf in _flatten(outputs): + if isinstance(leaf, Packed): + values[path] = leaf.layout.unpack(to_numpy(leaf.tensor)) + else: + values[path] = to_numpy(leaf) + self._free_transient(outputs) + pingpong = {s.name for s in self._slots.values() if s.pingpong} + if pingpong and (ctx.writes & pingpong) == pingpong: + self._phase ^= 1 + if isinstance(outputs, Packed) or not isinstance(outputs, (Mapping, list, tuple)): + return values[()] + return _rebuild(outputs, values) + + # ------------------------------------------------------------------------------------------- state + @property + def phase(self) -> int: + """Which buffer of each ping-pong state the next run reads (0 = the first buffer).""" + return self._phase + + def state_buffer(self, name: str): + """The device buffer the next run reads for state ``name`` (ping-pong: the buffer of the current + :attr:`phase`).""" + slot = self._slot(name, "state") + return slot.buffers[self._phase] if slot.pingpong else slot.buffers[0] + + def reset_state(self, name: Optional[str] = None, value: Any = None) -> None: + """Write the initial value (or ``value``) of state ``name`` (all states when ``None``) into the buffer the + next run reads. ``value``: numpy / scalar / torch / ttnn host tensor (uploaded) or a ttnn *device* tensor of + the same shape (``ttnn.copy`` on the device). Enqueued on CQ0, so it lands after any replay already + enqueued.""" + import ttnn + + self._check_open() + names = [name] if name is not None else [s.name for s in self._slots.values() if s.kind == "state"] + for n in names: + slot = self._slot(n, "state") + target = self.state_buffer(n) + if isinstance(value, ttnn.Tensor) and value.storage_type() == ttnn.StorageType.DEVICE: + _check_copy(f"reset_state({n!r})", value, slot) + with self._eager(f"reset_state({n!r}) from a device tensor"): + ttnn.copy(value, target) + continue + if value is None: + host = slot.init + elif isinstance(value, ttnn.Tensor): + host = to_host_tensor(value, slot.dtype, slot.layout, shape=slot.shape) + else: + arr = to_numpy(value) if hasattr(value, "detach") else np.asarray(value) + host = to_host_tensor(np.broadcast_to(arr, slot.shape), slot.dtype, slot.layout) + ttnn.copy_host_to_device_tensor(host, target, cq_id=CQ_COMPUTE) + + def read_state(self, name: str) -> np.ndarray: + """The value the next run will read for state ``name`` (blocking read on CQ0).""" + return to_numpy(self.state_buffer(name)) + + def _bank(self, name: str, bank: int): + slot = self._slot(name, "state") + if not 0 <= bank < len(slot.banks): + raise IndexError(f"{self.name}: state {name!r} has {len(slot.banks)} bank(s) (add_state(banks=...)), " + f"not bank {bank}") + return slot.banks[bank] + + def save_state(self, name: str, bank: int) -> None: + """Copy the current value of state ``name`` into its ``bank`` (``ttnn.copy`` on CQ0, eager, after any + replay already enqueued): the first half of a stream switch (PLAN.md D16; :class:`StreamBanks`).""" + import ttnn + + self._check_open() + with self._eager(f"save_state({name!r})"): + ttnn.copy(self.state_buffer(name), self._bank(name, bank)) + + def load_state(self, name: str, bank: int) -> None: + """Copy ``bank`` back into the buffer the next run reads for state ``name`` (``ttnn.copy`` on CQ0).""" + import ttnn + + self._check_open() + with self._eager(f"load_state({name!r})"): + ttnn.copy(self._bank(name, bank), self.state_buffer(name)) + + # ------------------------------------------------------------------------------------------- misc + def trace_ids(self) -> Dict[Tuple[str, int], Any]: + """``{(variant, phase): trace_id}`` for profiling scripts.""" + return {k: t.trace_id for k, t in self._traces.items()} + + def describe(self) -> Dict[str, Any]: + """A JSON-able summary for ``model.info`` / OPT_BASELINE.""" + def spec(s: _Slot) -> Dict[str, Any]: + return {"shape": list(s.shape), "dtype": dtype_name(s.dtype), "layout": str(s.layout).rsplit(".", 1)[-1], + **({"pingpong": s.pingpong, "banks": len(s.banks)} if s.kind == "state" else {})} + + entries = None + if hasattr(self.device, "num_program_cache_entries"): + entries = int(self.device.num_program_cache_entries()) + return { + "name": self.name, "num_command_queues": self.num_command_queues, "stage_inputs": self.stage_inputs, + "warmup_runs": self.warmup_runs, "phases": self.phases, "alloc_tracking": self.alloc_tracking, + "variants": sorted(self._variants), "traces": len(self._traces), + "inputs": {s.name: spec(s) for s in self._slots.values() if s.kind == "input"}, + "params": {s.name: spec(s) for s in self._slots.values() if s.kind == "param"}, + "states": {s.name: spec(s) for s in self._slots.values() if s.kind == "state"}, + "outputs": {s.name: spec(s) for s in self._slots.values() if s.kind == "output"}, + "timings_ms": {k: {kk: round(vv, 3) for kk, vv in v.items()} for k, v in self.timings_ms.items()}, + "program_cache_entries": entries, + } + + def _release_traces(self) -> None: + traces, self._traces = self._traces, {} + for trace in traces.values(): + self._safe_release(trace.trace_id) + self._last_phase.clear() + self._last_variant = None + self._captures = 0 + self._deallocate_outputs(tensor for trace in traces.values() for _, tensor, _ in trace.leaves) + + def release(self) -> None: + """Release every trace and deallocate the persistent tensors. Idempotent; the persistent tensors are freed + even when the device sync or a trace release raises (the error propagates afterwards).""" + import ttnn + + if self._closed: + return + self._closed = True + try: + try: + ttnn.synchronize_device(self.device) + finally: + self._release_traces() + finally: + slots, self._slots = list(self._slots.values()), {} + for slot in slots: + for t in slot.tensors(): + if t.is_allocated(): + ttnn.deallocate(t) + + def _check_open(self) -> None: + if self._closed: + raise RuntimeError(f"{self.name}: the TraceRunner was released") + + def __enter__(self) -> "TraceRunner": + return self + + def __exit__(self, *exc) -> None: + self.release() + + +class StreamBanks: + """Per-stream device state of a temporal model (PLAN.md D16) on top of the states of a :class:`TraceRunner`. + + The traces read and write one set of state buffers: the *active* stream's. With ``max_streams > 1`` every known + stream id owns one bank of each state (``add_state(..., banks=max_streams)``) and a switch saves the active + stream into its bank and loads the selected one (``ttnn.copy`` on CQ0, after the replays already enqueued; + probe P13 measured ~80 us eager for a StreamPETR-size state). A new id beyond ``max_streams`` is refused + (``on_full="reject"``: :class:`~.io.InputError`, HTTP 400) or takes over the least recently used stream + (``"evict"``; with ``max_streams=1`` that simply restarts the one state). A stream starts fresh -- its states + reset to their ``init`` values -- when it is new, when ``reset=True``, or when its timestamp goes backwards or + jumps by more than ``max_gap_s``; :meth:`select` returns True then, so the model can also set its first-frame + RT-dev params (BEVFormer ``use_prev_bev=0``, BEVDet ``flag``, ...). Create it after the last ``capture()`` (a + capture resets every state) and call :meth:`select` under the model lock, before the run of each frame:: + + S_MAX = 1 # first publish (D16) + runner.add_state("prev_bev", shape=(1, 1, 22500, 256), dtype="bfloat16", pingpong=True, + banks=S_MAX if S_MAX > 1 else 0) + streams = StreamBanks(runner, ["prev_bev"], max_streams=S_MAX, on_full="evict", max_gap_s=2.0) + ... + fresh = streams.select(stream.get("id", "default"), reset=stream.get("reset", False), + timestamp_s=stream.get("timestamp_s")) + out = runner("frame", inputs=..., params={"use_prev_bev": 0.0 if fresh else 1.0}) + """ + + def __init__(self, runner: TraceRunner, states: Sequence[str], *, max_streams: int = 1, on_full: str = "reject", + max_gap_s: Optional[float] = None): + if int(max_streams) < 1: + raise ValueError("max_streams must be >= 1") + if on_full not in ("reject", "evict"): + raise ValueError(f"on_full={on_full!r}: expected 'reject' or 'evict'") + self.runner = runner + self.states = tuple(states) + if not self.states: + raise ValueError("StreamBanks needs at least one state") + self.max_streams = int(max_streams) + for name in self.states: + banks = len(runner._slot(name, "state").banks) + if self.max_streams > 1 and banks < self.max_streams: + raise ValueError(f"state {name!r} has {banks} bank(s); StreamBanks(max_streams={self.max_streams}) " + f"needs add_state(..., banks={self.max_streams})") + self.on_full = on_full + self.max_gap_s = None if max_gap_s is None else float(max_gap_s) + self.active: Optional[str] = None + self._bank: Dict[str, int] = {} # stream id -> bank index (max_streams > 1) + self._last_t: Dict[str, float] = {} # stream id -> timestamp of its last frame + self._used: Dict[str, int] = {} # stream id -> last use (LRU clock) + self._clock = 0 + + @property + def streams(self) -> List[str]: + """Known stream ids, most recently used first.""" + return sorted(self._used, key=lambda s: -self._used[s]) + + def _save_active(self) -> None: + if self.active is not None and self.active in self._bank: + for name in self.states: + self.runner.save_state(name, self._bank[self.active]) + + def forget(self, stream_id: str) -> None: + """Drop a stream (its bank becomes free; if it was active, the next :meth:`select` starts fresh).""" + sid = str(stream_id) + self._used.pop(sid, None) + self._bank.pop(sid, None) + self._last_t.pop(sid, None) + if self.active == sid: + self.active = None + + def select(self, stream_id: Any = "default", *, reset: bool = False, timestamp_s: Optional[float] = None) -> bool: + """Make ``stream_id`` the active stream for the next run; returns True when its state starts fresh.""" + sid = str(stream_id) + fresh = bool(reset) + if sid != self.active: + if sid in self._used: # a known stream parked in its bank + self._save_active() + if not fresh: + for name in self.states: + self.runner.load_state(name, self._bank[sid]) + else: # a new stream + if len(self._used) >= self.max_streams: + if self.on_full == "reject": + raise InputError(f"stream {sid!r}: this model keeps device state for {self.max_streams} " + f"stream(s), in use by {self.streams}; reuse an id") + self.forget(min(self._used, key=self._used.__getitem__)) + self._save_active() + if self.max_streams > 1: + self._bank[sid] = min(set(range(self.max_streams)) - set(self._bank.values())) + fresh = True + self.active = sid + if not fresh and timestamp_s is not None and self.max_gap_s is not None and sid in self._last_t: + dt = float(timestamp_s) - self._last_t[sid] + fresh = dt < 0 or dt > self.max_gap_s + if fresh: + for name in self.states: + self.runner.reset_state(name) + if timestamp_s is not None: + self._last_t[sid] = float(timestamp_s) + elif fresh: + self._last_t.pop(sid, None) + self._clock += 1 + self._used[sid] = self._clock + return fresh + + def describe(self) -> Dict[str, Any]: + """JSON-able summary for ``model.info``.""" + return {"max_streams": self.max_streams, "on_full": self.on_full, "max_gap_s": self.max_gap_s, + "states": list(self.states), "active": self.active, "streams": self.streams} diff --git a/code/tt_diffusion_planner/ttaw/voxelize.py b/code/tt_diffusion_planner/ttaw/voxelize.py new file mode 100644 index 0000000000000000000000000000000000000000..755df00522b42786343b105b06cce7f76defea92 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/voxelize.py @@ -0,0 +1,260 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C12 voxelization: deterministic pillar builders (vectorised numpy), Autoware's pillar feature decoration, and the +two canvas forms of the PointPillars scatter (dense scatter for references, gather index for the device). + +Autoware's CUDA voxelizer (``autoware_lidar_centerpoint/lib/preprocess/preprocess_kernel.cu``) is +non-deterministic: slots inside a pillar come from ``atomicAdd`` over a randomly shuffled point buffer and pillar +ids from ``atomicAdd`` over a 2-D thread grid (S:centerpoint:177-186). The ports replace it by a deterministic +policy with the same statistics (PLAN.md D1, S:centerpoint:381-389): + +1. optional seeded permutation of the input rows (:func:`ttaw.pointcloud.seeded_permutation`); +2. the half-open range test and the cell index ``floor((v - min) / voxel)`` (a float32 *division*, clamped; the + kernel's ``:135-145``); +3. the first ``max_points_per_pillar`` points of every cell in (permuted) input order; +4. pillar ids in ascending cell order, keeping the first ``max_pillars``. The cell order is ``"flipped_x"`` + (CenterPoint: ``cell = (grid_x - 1 - ix) * grid_y + iy``, front rows first, ``:144``) or ``"raster"`` + (PointPainting: ``cell = iy * grid_x + ix``, S:pointpainting:229-232). + +The result (:class:`PillarSet`) holds the raw rows of every kept slot (zero-padded), the point counts and the +``(z, y, x)`` = ``(0, iy, ix)`` coordinates of mmdet3d / the CUDA ``uint3 {0, idy, idx}`` (``:189-190``). + +:func:`decorate_pillar_features` is ``generateFeatures_kernel`` (``:211-343``) for 9 / 10 / 11 encoder features +(the cluster mean is summed slot by slot in float32, a fixed order standing in for the kernel's float atomics; the +pillar centre is ``voxel / 2 + coord * voxel + min`` in float32; padded slots are all zero). + +Canvas: :func:`scatter_canvas` is ``scatterFeatures_kernel`` (``scatter_kernel.cu:25-48``: ``canvas[c, iy, ix]``). +The device builds the same canvas as a *gather* (C19, probe P4): :func:`canvas_gather_index` gives, for every cell +``iy * grid_x + ix`` of the NHWC canvas, the row of its pillar or a sentinel row (an explicit zero row of the +feature table), so the canvas is ``table[index]`` with no scatter, no clear and a static shape. + +numpy only; no side effects on import. +""" +from __future__ import annotations + +from dataclasses import dataclass +from typing import Any, Dict, Optional, Sequence, Tuple + +import numpy as np + +from .pointcloud import seeded_permutation + +__all__ = [ + "PillarGridSpec", + "PillarSet", + "assign_pillars", + "decorate_pillar_features", + "scatter_canvas", + "canvas_gather_index", + "pillar_cell_index", +] + +_ORDERS = ("flipped_x", "raster") + + +@dataclass(frozen=True) +class PillarGridSpec: + """A pillar grid and its capacities. Ranges and sizes are float32, as the CUDA config stores them + (``centerpoint_config.hpp:41-56``); ``grid_x = int((max_x - min_x) / voxel_x)`` in float32 (``:95-97``).""" + + range_min: Tuple[float, float, float] + range_max: Tuple[float, float, float] + voxel_size: Tuple[float, float, float] + max_points_per_pillar: int = 32 + max_pillars: int = 40000 + order: str = "flipped_x" + + def __post_init__(self) -> None: + if self.order not in _ORDERS: + raise ValueError(f"order must be one of {_ORDERS}, got {self.order!r}") + if self.max_points_per_pillar < 1 or self.max_pillars < 1: + raise ValueError("capacities must be >= 1") + + @property + def min_f32(self) -> np.ndarray: + return np.asarray(self.range_min, dtype=np.float32) + + @property + def max_f32(self) -> np.ndarray: + return np.asarray(self.range_max, dtype=np.float32) + + @property + def voxel_f32(self) -> np.ndarray: + return np.asarray(self.voxel_size, dtype=np.float32) + + def _grid(self, axis: int) -> int: + return int(np.float32(self.max_f32[axis] - self.min_f32[axis]) / self.voxel_f32[axis]) + + @property + def grid_x(self) -> int: + return self._grid(0) + + @property + def grid_y(self) -> int: + return self._grid(1) + + @property + def grid_z(self) -> int: + return self._grid(2) + + @property + def num_cells(self) -> int: + return self.grid_x * self.grid_y + + +@dataclass +class PillarSet: + """The deterministic voxelization of one cloud (fixed capacity ``max_pillars``; rows past ``num_pillars`` are + zero).""" + + points: np.ndarray # (max_pillars, K, F) float32: raw rows of the kept slots, zero-padded + num_points: np.ndarray # (max_pillars,) int32: min(count, K), 0 for unused rows + coords: np.ndarray # (max_pillars, 3) int32: (0, iy, ix) + num_pillars: int # kept pillars (<= max_pillars) + num_pillars_total: int # non-empty cells before the cap (Autoware's num_voxels counter) + num_points_in_range: int + max_points_in_pillar: int # largest per-cell count before the slot cap + pillars_over_cap: int # cells holding more than K points + num_points_input: int + + @property + def overflow(self) -> bool: + """More non-empty cells than ``max_pillars``: the farthest-back pillars were dropped.""" + return self.num_pillars_total > self.points.shape[0] + + def stats(self) -> Dict[str, Any]: + return {"num_points": self.num_points_input, "num_points_in_range": self.num_points_in_range, + "num_pillars": self.num_pillars, "num_pillars_total": self.num_pillars_total, + "max_points_in_pillar": self.max_points_in_pillar, "pillars_over_cap": self.pillars_over_cap, + "pillar_overflow": self.overflow} + + +def pillar_cell_index(points: np.ndarray, spec: PillarGridSpec) -> Tuple[np.ndarray, np.ndarray]: + """``(ix, iy)`` int64 of each row: ``floor((v - min) / voxel)`` evaluated in float32, clamped to the grid + (``preprocess_kernel.cu:140-143``). The caller has already applied the range test.""" + p = np.asarray(points, dtype=np.float32) + mn, vs = spec.min_f32, spec.voxel_f32 + ix = np.floor((p[:, 0] - mn[0]) / vs[0]).astype(np.int64) + iy = np.floor((p[:, 1] - mn[1]) / vs[1]).astype(np.int64) + return np.clip(ix, 0, spec.grid_x - 1), np.clip(iy, 0, spec.grid_y - 1) + + +def _cell_key(ix: np.ndarray, iy: np.ndarray, spec: PillarGridSpec) -> np.ndarray: + if spec.order == "flipped_x": + return (spec.grid_x - 1 - ix) * spec.grid_y + iy + return iy * spec.grid_x + ix + + +def _key_to_ixy(key: np.ndarray, spec: PillarGridSpec) -> Tuple[np.ndarray, np.ndarray]: + if spec.order == "flipped_x": + return (spec.grid_x - 1) - key // spec.grid_y, key % spec.grid_y + return key % spec.grid_x, key // spec.grid_x + + +def assign_pillars(points: np.ndarray, spec: PillarGridSpec, *, shuffle_seed: Optional[int] = None) -> PillarSet: + """Deterministic voxelization (module docstring). ``points``: (N, F) float32 in the model frame, x, y, z in + columns 0-2 (the remaining columns, e.g. time_lag, are copied into the slots). Rows with a non-finite x, y or z + never enter a pillar.""" + pts = np.ascontiguousarray(np.asarray(points, dtype=np.float32)) + if pts.ndim != 2 or pts.shape[1] < 3: + raise ValueError(f"points must be (N, >=3), got {pts.shape}") + n_in = len(pts) + perm = seeded_permutation(n_in, shuffle_seed) + if perm is not None: + pts = pts[perm] + mn, mx = spec.min_f32, spec.max_f32 + keep = np.ones(len(pts), dtype=bool) + for a in range(3): # Autoware: reject x < min || x >= max (NaN would pass there; here it fails) + keep &= (pts[:, a] >= mn[a]) & (pts[:, a] < mx[a]) + pts = pts[keep] + ix, iy = pillar_cell_index(pts, spec) + key = _cell_key(ix, iy, spec) + order = np.argsort(key, kind="stable") # stable: input order inside each cell + key_sorted = key[order] + uniq, first, counts = np.unique(key_sorted, return_index=True, return_counts=True) + slot = np.arange(len(key_sorted)) - np.repeat(first, counts) + K, P = spec.max_points_per_pillar, spec.max_pillars + n_total = int(len(uniq)) + n_keep = min(n_total, P) + pid = np.repeat(np.arange(n_total), counts) # pillar id of every sorted row (ascending cell order) + sel = (slot < K) & (pid < n_keep) + F = pts.shape[1] + raw = np.zeros((P, K, F), dtype=np.float32) + raw[pid[sel], slot[sel]] = pts[order][sel] + num = np.zeros(P, dtype=np.int32) + num[:n_keep] = np.minimum(counts[:n_keep], K) + cx, cy = _key_to_ixy(uniq[:n_keep], spec) + coords = np.zeros((P, 3), dtype=np.int32) + coords[:n_keep, 1] = cy + coords[:n_keep, 2] = cx + return PillarSet(points=raw, num_points=num, coords=coords, num_pillars=n_keep, num_pillars_total=n_total, + num_points_in_range=int(keep.sum()), max_points_in_pillar=int(counts.max()) if n_total else 0, + pillars_over_cap=int((counts > K).sum()), num_points_input=n_in) + + +def decorate_pillar_features(pillars: PillarSet, spec: PillarGridSpec, *, encoder_in_feature_size: int = 9, + point_dim: Optional[int] = None) -> np.ndarray: + """``generateFeatures_kernel`` (``preprocess_kernel.cu:211-343``) -> (max_pillars, K, encoder_in_feature_size) + float32. ``point_dim`` (columns copied from the raw rows) is 5 (x, y, z, intensity, time_lag) for 11 features + and 4 (x, y, z, time_lag) otherwise, as the kernel's ``point_dim``. Layouts: + + - 9: x, y, z, t, x-x̄, y-ȳ, z-z̄, x-x_c, y-y_c + - 10: x, y, z, t, x-x̄, y-ȳ, z-z̄, x-x_c, y-y_c, z-z_c + - 11: x, y, z, i, t, x-x̄, y-ȳ, z-z̄, x-x_c, y-y_c, z-z_c + + x̄ is the mean of the pillar's kept points (summed slot by slot in float32, divided by the float count); + ``x_c = voxel_x / 2 + ix * voxel_x + min_x`` in float32 (``:269-276``). Padded slots and unused pillars are 0.""" + F_enc = int(encoder_in_feature_size) + if F_enc not in (9, 10, 11): + raise ValueError("encoder_in_feature_size must be 9, 10 or 11") + pd = (5 if F_enc >= 11 else 4) if point_dim is None else int(point_dim) + raw = pillars.points + P, K, F = raw.shape + if F < pd: + raise ValueError(f"pillar rows have {F} columns, decoration needs {pd}") + n = pillars.num_pillars + out = np.zeros((P, K, F_enc), dtype=np.float32) + if n == 0: + return out + r = raw[:n] + cnt = pillars.num_points[:n].astype(np.float32) + s = np.zeros((n, 3), dtype=np.float32) + for k in range(K): # fixed slot order (the kernel's float atomicAdd order is unspecified) + s += r[:, k, :3] + mean = s / cnt[:, None] + vs, mn = spec.voxel_f32, spec.min_f32 + half = vs / np.float32(2) + xo = (half[0] + pillars.coords[:n, 2].astype(np.float32) * vs[0]) + mn[0] + yo = (half[1] + pillars.coords[:n, 1].astype(np.float32) * vs[1]) + mn[1] + zo = (half[2] + pillars.coords[:n, 0].astype(np.float32) * vs[2]) + mn[2] + f = out[:n] + f[:, :, :pd] = r[:, :, :pd] + f[:, :, pd:pd + 3] = r[:, :, :3] - mean[:, None, :] + f[:, :, pd + 3] = r[:, :, 0] - xo[:, None] + f[:, :, pd + 4] = r[:, :, 1] - yo[:, None] + if F_enc in (10, 11): + f[:, :, pd + 5] = r[:, :, 2] - zo[:, None] + valid = np.arange(K)[None, :] < pillars.num_points[:n, None] + f[~valid] = 0.0 + return out + + +def scatter_canvas(features: np.ndarray, coords: np.ndarray, num_pillars: int, spec: PillarGridSpec) -> np.ndarray: + """``scatterFeatures_kernel``: ``canvas[c, iy, ix] = features[p, c]`` for the first ``num_pillars`` rows; the + canvas is (C, grid_y, grid_x) float32 and zero elsewhere (NCHW without the batch axis).""" + f = np.asarray(features, dtype=np.float32).reshape(np.shape(features)[0], -1) + C = f.shape[1] + canvas = np.zeros((C, spec.grid_y, spec.grid_x), dtype=np.float32) + c = np.asarray(coords)[:num_pillars] + canvas[:, c[:, 1], c[:, 2]] = f[:num_pillars].T + return canvas + + +def canvas_gather_index(coords: np.ndarray, num_pillars: int, spec: PillarGridSpec, *, sentinel: int, + dtype: Any = np.uint32) -> np.ndarray: + """Gather form of the scatter (C19): for every cell ``iy * grid_x + ix`` of the row-major (NHWC) canvas, the row + of its pillar, or ``sentinel`` (an all-zero row of the feature table) for an empty cell. Then + ``canvas_nhwc = table[index]`` equals :func:`scatter_canvas` transposed to (grid_y, grid_x, C).""" + idx = np.full(spec.num_cells, sentinel, dtype=np.int64) + c = np.asarray(coords)[:num_pillars] + idx[c[:, 1].astype(np.int64) * spec.grid_x + c[:, 2].astype(np.int64)] = np.arange(num_pillars) + return idx.astype(dtype) diff --git a/code/tt_diffusion_planner/ttaw/weights.py b/code/tt_diffusion_planner/ttaw/weights.py new file mode 100644 index 0000000000000000000000000000000000000000..4f7dd87a2d28408db9bb3d8b0cf377343b2bcd03 --- /dev/null +++ b/code/tt_diffusion_planner/ttaw/weights.py @@ -0,0 +1,632 @@ +# SPDX-License-Identifier: Apache-2.0 +"""C04: weights -- ONNX initializers addressed by the node that consumes them, BN folding in float64, safetensors / +``.pth`` loading, and an on-disk cache of prepared (tilized, converted) device tensors. + +Why by consuming node: initializer names drift between exports of the same network (FRNet), and MatMul weights are +often anonymous (``onnx::MatMul_1234`` in Diffusion Planner), while node names (``/backbone/conv1/Conv``) are stable. +``OnnxWeights`` parses the graph as data only (no onnxruntime, no execution); external-data paths must stay inside +the model's directory. + +BN folding: every fold is computed in float64 and rounded once at the end (default float32; pass ``dtype=None`` to +keep float64 and round straight to the device dtype later). + +The device cache stores ``ttnn`` host tensors (``.tensorbin``) under ``$TT_CACHE_PATH`` (the tt-model container +mounts it at ``/tensor-cache``) or ``~/.cache/ttaw``, keyed by a digest of the source weights, so a new weights +revision never reuses stale tensors. ``TTAW_WEIGHT_CACHE=0`` disables it (A/B switch, read when the cache is built). + +Only ``numpy`` is imported at module level; ``onnx``, ``safetensors``, ``torch`` and ``ttnn`` are imported by the +functions that need them. +""" +from __future__ import annotations + +import fnmatch +import hashlib +import os +import re +from dataclasses import dataclass, field +from pathlib import Path +from typing import Any, Callable, Dict, Iterator, List, Mapping, Optional, Tuple, Union + +import numpy as np + +__all__ = [ + "BatchNorm", + "fold_bn", + "fold_bn_conv", + "fold_bn_linear", + "fold_bn_conv_transpose", + "OnnxNode", + "ConvParams", + "GemmParams", + "OnnxWeights", + "StateDict", + "load_safetensors", + "load_torch_checkpoint", + "load_state_dict", + "file_sha256", + "cache_root", + "WeightCache", +] + +PathLike = Union[str, os.PathLike] + + +# ---------------------------------------------------------------------------------------------- BN folding + +@dataclass(frozen=True) +class BatchNorm: + """Inference BatchNorm parameters (any of BN1d / BN2d / ONNX BatchNormalization), stored in float64.""" + + gamma: np.ndarray + beta: np.ndarray + mean: np.ndarray + var: np.ndarray + eps: float = 1e-5 + + def __post_init__(self) -> None: + for name in ("gamma", "beta", "mean", "var"): + object.__setattr__(self, name, np.asarray(getattr(self, name), dtype=np.float64).reshape(-1)) + n = self.gamma.shape[0] + if not (self.beta.shape[0] == self.mean.shape[0] == self.var.shape[0] == n): + raise ValueError("BatchNorm parameter lengths differ") + + @property + def channels(self) -> int: + return int(self.gamma.shape[0]) + + @property + def scale(self) -> np.ndarray: + """``gamma / sqrt(var + eps)`` (float64).""" + return self.gamma / np.sqrt(self.var + self.eps) + + @property + def shift(self) -> np.ndarray: + """``beta - mean * scale`` (float64): ``bn(x) == x * scale + shift``.""" + return self.beta - self.mean * self.scale + + @classmethod + def from_state_dict(cls, sd: Mapping[str, Any], prefix: str, eps: float = 1e-5) -> "BatchNorm": + """torch naming: ``.weight``, ``.bias``, ``.running_mean``, ``.running_var``.""" + p = prefix.rstrip(".") + "." if prefix else "" + return cls(np.asarray(sd[p + "weight"]), np.asarray(sd[p + "bias"]), np.asarray(sd[p + "running_mean"]), + np.asarray(sd[p + "running_var"]), eps) + + def apply(self, x: np.ndarray, axis: int = 1) -> np.ndarray: + """Reference application in float64 (``axis`` = channel axis), for tests.""" + shape = [1] * np.ndim(x) + shape[axis] = -1 + return np.asarray(x, np.float64) * self.scale.reshape(shape) + self.shift.reshape(shape) + + +def _finish(w: np.ndarray, b: np.ndarray, dtype: Optional[Any]) -> Tuple[np.ndarray, np.ndarray]: + if dtype is None: + return w, b + return w.astype(dtype), b.astype(dtype) + + +def fold_bn(weight: np.ndarray, bias: Optional[np.ndarray], bn: BatchNorm, *, axis: int = 0, + dtype: Optional[Any] = np.float32) -> Tuple[np.ndarray, np.ndarray]: + """Fold ``bn`` (applied to the layer's output) into ``weight`` / ``bias``: ``W' = W * s`` along the output axis + ``axis`` and ``b' = (b - mean) * s + beta``, all in float64, one final rounding to ``dtype``.""" + w = np.asarray(weight, np.float64) + if w.shape[axis] != bn.channels: + raise ValueError(f"weight axis {axis} has {w.shape[axis]} channels, BatchNorm has {bn.channels}") + shape = [1] * w.ndim + shape[axis] = -1 + s = bn.scale + b = np.zeros(bn.channels, np.float64) if bias is None else np.asarray(bias, np.float64).reshape(-1) + return _finish(w * s.reshape(shape), (b - bn.mean) * s + bn.beta, dtype) + + +def fold_bn_conv(weight: np.ndarray, bias: Optional[np.ndarray], bn: BatchNorm, *, + dtype: Optional[Any] = np.float32) -> Tuple[np.ndarray, np.ndarray]: + """Conv1d/2d/3d weight ``[Cout, Cin/groups, k...]`` followed by BN (any ``groups``).""" + return fold_bn(weight, bias, bn, axis=0, dtype=dtype) + + +def fold_bn_linear(weight: np.ndarray, bias: Optional[np.ndarray], bn: BatchNorm, *, layout: str = "out_in", + dtype: Optional[Any] = np.float32) -> Tuple[np.ndarray, np.ndarray]: + """Linear / Gemm / MatMul followed by BN1d. ``layout="out_in"``: torch ``Linear`` and ONNX Gemm with + ``transB=1`` (``[out, in]``); ``"in_out"``: ONNX MatMul / Gemm ``transB=0`` (``[in, out]``).""" + if layout not in ("out_in", "in_out"): + raise ValueError("layout must be 'out_in' or 'in_out'") + return fold_bn(weight, bias, bn, axis=0 if layout == "out_in" else -1, dtype=dtype) + + +def fold_bn_conv_transpose(weight: np.ndarray, bias: Optional[np.ndarray], bn: BatchNorm, *, groups: int = 1, + dtype: Optional[Any] = np.float32) -> Tuple[np.ndarray, np.ndarray]: + """ConvTranspose weight ``[Cin, Cout/groups, k...]`` followed by BN: output channel of ``W[i, j]`` is + ``(i // (Cin/groups)) * (Cout/groups) + j``.""" + w = np.asarray(weight, np.float64) + cin, cout_g = w.shape[0], w.shape[1] + if cin % groups: + raise ValueError(f"Cin={cin} is not divisible by groups={groups}") + if cout_g * groups != bn.channels: + raise ValueError(f"ConvTranspose has {cout_g * groups} output channels, BatchNorm has {bn.channels}") + out_idx = (np.arange(cin)[:, None] // (cin // groups)) * cout_g + np.arange(cout_g)[None, :] + s = bn.scale[out_idx].reshape((cin, cout_g) + (1,) * (w.ndim - 2)) + b = np.zeros(bn.channels, np.float64) if bias is None else np.asarray(bias, np.float64).reshape(-1) + return _finish(w * s, (b - bn.mean) * bn.scale + bn.beta, dtype) + + +# ----------------------------------------------------------------------------------------------- ONNX + +@dataclass(frozen=True) +class OnnxNode: + """A graph node: name, op type, input / output tensor names and decoded attributes.""" + + name: str + op_type: str + inputs: Tuple[str, ...] + outputs: Tuple[str, ...] + attrs: Dict[str, Any] = field(default_factory=dict) + index: int = 0 + + +@dataclass(frozen=True) +class ConvParams: + """Weights and attributes of a Conv / ConvTranspose node. ``pads`` is the explicit attribute (zeros when + absent); with ``auto_pad`` ``SAME_UPPER`` / ``SAME_LOWER`` the real padding depends on the input size and must + be derived from it (PyTorch exports always write explicit pads and ``NOTSET``).""" + + weight: np.ndarray + bias: Optional[np.ndarray] + strides: Tuple[int, ...] + pads: Tuple[int, ...] + dilations: Tuple[int, ...] + group: int + kernel_shape: Tuple[int, ...] + output_padding: Tuple[int, ...] = () + auto_pad: str = "NOTSET" + output_shape: Tuple[int, ...] = () + + +@dataclass(frozen=True) +class GemmParams: + """Weights and attributes of a Gemm node (``weight`` exactly as stored; see ``trans_b``).""" + + weight: np.ndarray + bias: Optional[np.ndarray] + trans_a: bool + trans_b: bool + alpha: float + beta: float + + +def _attr_value(attr) -> Any: + from onnx import helper + + value = helper.get_attribute_value(attr) + if isinstance(value, bytes): + return value.decode("utf-8", "replace") + if isinstance(value, list) and value and isinstance(value[0], bytes): + return [v.decode("utf-8", "replace") for v in value] + return value + + +class OnnxWeights: + """Initializers and Constant tensors of an ONNX file, addressable by tensor name or by consuming node. + + Example:: + + w = OnnxWeights("pts_backbone_neck_head_centerpoint.onnx") + conv = w.conv("/backbone/blocks.0/blocks.0.0/Conv") # ConvParams(weight, bias, strides, ...) + wf, bf = w.fold_conv_bn("/backbone/blocks.0/blocks.0.0/Conv") # conv + its BN consumer, fp64 fold + k = w.param("/dit/blocks.0/attn/MatMul", 1) # anonymous MatMul weight by slot + for node in w.nodes("Conv", "/backbone/*"): # graph order, glob on node names + print(node.name, w.conv(node.name).weight.shape) + """ + + def __init__(self, path: PathLike, *, load_external_data: bool = True): + import onnx + + self.path = Path(path) + model = onnx.load(str(self.path), load_external_data=False) + if load_external_data: + self._load_external(model) + self._sha256: Optional[str] = None + graph = model.graph + self._tensors: Dict[str, Any] = {t.name: t for t in graph.initializer} + self._cache: Dict[str, np.ndarray] = {} + self._nodes: Dict[str, OnnxNode] = {} + self._order: List[str] = [] + self._producer: Dict[str, str] = {} + self._consumers: Dict[str, List[str]] = {} + self._alias: Dict[str, str] = {} + for i, n in enumerate(graph.node): + name = n.name or f"{n.op_type}_{i}" + if name in self._nodes: + name = f"{name}#{i}" + node = OnnxNode(name, n.op_type, tuple(n.input), tuple(n.output), + {a.name: _attr_value(a) for a in n.attribute}, i) + self._nodes[name] = node + self._order.append(name) + for out in n.output: + self._producer[out] = name + for inp in n.input: + if inp: + self._consumers.setdefault(inp, []).append(name) + if n.op_type == "Constant" and n.output: + self._add_constant(n.output[0], n) + for name in self._order: # Identity of a constant tensor is that tensor + node = self._nodes[name] + if node.op_type == "Identity" and node.inputs and node.outputs: + self._alias[node.outputs[0]] = node.inputs[0] + self.input_names = tuple(i.name for i in graph.input if i.name not in self._tensors) + self.output_names = tuple(o.name for o in graph.output) + self.opset = {o.domain or "ai.onnx": int(o.version) for o in model.opset_import} + + def _load_external(self, model) -> None: + from onnx import TensorProto + from onnx import external_data_helper as ext + + base = self.path.parent.resolve() + for tensor in model.graph.initializer: + if tensor.data_location != TensorProto.EXTERNAL: + continue + location = {e.key: e.value for e in tensor.external_data}.get("location", "") + target = (base / location).resolve() + if base not in target.parents: + raise ValueError(f"external data of {tensor.name!r} points outside {base}: {location!r}") + ext.load_external_data_for_tensor(tensor, str(base)) + tensor.data_location = TensorProto.DEFAULT + del tensor.external_data[:] + + def _add_constant(self, output: str, node) -> None: + from onnx import numpy_helper + + for attr in node.attribute: + if attr.name == "value": + self._tensors[output] = attr.t + elif attr.name in ("value_float", "value_floats", "value_int", "value_ints"): + dt = np.float32 if "float" in attr.name else np.int64 + self._cache[output] = np.asarray(_attr_value(attr), dtype=dt) + self._tensors[output] = numpy_helper.from_array(self._cache[output], output) + + # ---- tensors --------------------------------------------------------------------------------- + @property + def sha256(self) -> str: + """sha256 of the ``.onnx`` file (computed on first use; a cache-key ingredient).""" + if self._sha256 is None: + self._sha256 = file_sha256(self.path) + return self._sha256 + + def _resolve(self, name: str) -> str: + seen = set() + while name in self._alias and name not in self._tensors: + if name in seen: + raise ValueError(f"Identity cycle at {name!r}") + seen.add(name) + name = self._alias[name] + return name + + def has(self, name: str) -> bool: + """True if ``name`` is an initializer / Constant output (or an Identity of one).""" + return self._resolve(name) in self._tensors + + def array(self, name: str) -> np.ndarray: + """The constant tensor ``name`` as numpy (cached; treat it as read-only).""" + from onnx import numpy_helper + + key = self._resolve(name) + if key not in self._tensors: + raise KeyError(f"{name!r} is not an initializer or Constant output of {self.path.name}") + if key not in self._cache: + self._cache[key] = numpy_helper.to_array(self._tensors[key]) + return self._cache[key] + + def initializer_names(self) -> List[str]: + return sorted(self._tensors) + + def state_dict(self) -> Dict[str, np.ndarray]: + """Every constant tensor by name (the export's names; prefer node addressing for stability).""" + return {name: self.array(name) for name in self._tensors} + + # ---- nodes ----------------------------------------------------------------------------------- + def node(self, name: str) -> OnnxNode: + try: + return self._nodes[name] + except KeyError: + close = [n for n in self._order if name in n][:5] + raise KeyError(f"no node {name!r} in {self.path.name}" + (f"; similar: {close}" if close else "")) from None + + def nodes(self, op_type: Optional[str] = None, pattern: Optional[str] = None, *, + regex: bool = False) -> List[OnnxNode]: + """Nodes in graph order, filtered by op type and a glob (or ``regex=True``) on the node name.""" + rx = re.compile(pattern) if (pattern and regex) else None + out = [] + for name in self._order: + node = self._nodes[name] + if op_type is not None and node.op_type != op_type: + continue + if pattern is not None and not (rx.search(name) if rx else fnmatch.fnmatchcase(name, pattern)): + continue + out.append(node) + return out + + def find_node(self, pattern: str, op_type: Optional[str] = None, *, regex: bool = False) -> OnnxNode: + """Exactly one node matching ``pattern`` (and ``op_type``), else ``KeyError``.""" + found = self.nodes(op_type, pattern, regex=regex) + if len(found) != 1: + raise KeyError(f"{len(found)} nodes match {pattern!r} (op_type={op_type}): {[n.name for n in found[:8]]}") + return found[0] + + def producer(self, tensor: str) -> Optional[OnnxNode]: + name = self._producer.get(tensor) + return self._nodes[name] if name else None + + def consumers(self, tensor: str) -> List[OnnxNode]: + """Nodes reading tensor ``tensor`` (directly; Identity nodes are listed as consumers).""" + return [self._nodes[n] for n in self._consumers.get(tensor, [])] + + def consumer_of(self, tensor: str, op_type: Optional[str] = None) -> OnnxNode: + """The single consumer of ``tensor`` (optionally of ``op_type``), else ``KeyError``.""" + found = [n for n in self.consumers(tensor) if op_type is None or n.op_type == op_type] + if len(found) != 1: + raise KeyError(f"{tensor!r} has {len(found)} consumers of type {op_type}: {[n.name for n in found]}") + return found[0] + + def param(self, node_name: str, index: int) -> np.ndarray: + """The constant feeding input slot ``index`` of node ``node_name``.""" + node = self.node(node_name) + if index >= len(node.inputs) or not node.inputs[index]: + raise KeyError(f"node {node_name!r} ({node.op_type}) has no input {index}") + name = node.inputs[index] + if not self.has(name): + raise KeyError(f"input {index} of {node_name!r} is {name!r}, a runtime tensor, not a constant") + return self.array(name) + + def params(self, node_name: str) -> Dict[int, np.ndarray]: + """``{slot: array}`` for every constant input of the node.""" + node = self.node(node_name) + return {i: self.array(t) for i, t in enumerate(node.inputs) if t and self.has(t)} + + def conv(self, node_name: str) -> ConvParams: + """Conv / ConvTranspose weights (+ bias) and attributes with ONNX defaults filled in.""" + node = self.node(node_name) + if node.op_type not in ("Conv", "ConvTranspose"): + raise ValueError(f"{node_name!r} is {node.op_type}, not Conv/ConvTranspose") + w = self.param(node_name, 1) + b = self.param(node_name, 2) if len(node.inputs) > 2 and node.inputs[2] else None + k = tuple(node.attrs.get("kernel_shape", w.shape[2:])) + nd = len(k) + attrs = node.attrs + return ConvParams(w, b, tuple(attrs.get("strides", (1,) * nd)), tuple(attrs.get("pads", (0,) * 2 * nd)), + tuple(attrs.get("dilations", (1,) * nd)), int(attrs.get("group", 1)), k, + tuple(attrs.get("output_padding", ())), str(attrs.get("auto_pad", "NOTSET")), + tuple(attrs.get("output_shape", ()))) + + def gemm(self, node_name: str) -> GemmParams: + node = self.node(node_name) + if node.op_type != "Gemm": + raise ValueError(f"{node_name!r} is {node.op_type}, not Gemm") + b = self.param(node_name, 2) if len(node.inputs) > 2 and node.inputs[2] else None + return GemmParams(self.param(node_name, 1), b, bool(node.attrs.get("transA", 0)), + bool(node.attrs.get("transB", 0)), float(node.attrs.get("alpha", 1.0)), + float(node.attrs.get("beta", 1.0))) + + def matmul_weight(self, node_name: str) -> np.ndarray: + """The constant operand of a MatMul (``[in, out]`` when it is the right operand).""" + node = self.node(node_name) + consts = [i for i, t in enumerate(node.inputs[:2]) if t and self.has(t)] + if len(consts) != 1: + raise ValueError(f"MatMul {node_name!r} has {len(consts)} constant operands") + return self.array(node.inputs[consts[0]]) + + def batchnorm(self, node_name: str) -> BatchNorm: + node = self.node(node_name) + if node.op_type != "BatchNormalization": + raise ValueError(f"{node_name!r} is {node.op_type}, not BatchNormalization") + return BatchNorm(self.param(node_name, 1), self.param(node_name, 2), self.param(node_name, 3), + self.param(node_name, 4), float(node.attrs.get("epsilon", 1e-5))) + + def fold_conv_bn(self, conv_node: str, bn_node: Optional[str] = None, *, + dtype: Optional[Any] = np.float32) -> Tuple[np.ndarray, np.ndarray]: + """Conv / ConvTranspose weights with the following BatchNormalization folded in (float64, one rounding). + ``bn_node`` defaults to the single BatchNormalization consuming the conv output.""" + conv = self.conv(conv_node) + if bn_node is None: + bn_node = self.consumer_of(self.node(conv_node).outputs[0], "BatchNormalization").name + bn = self.batchnorm(bn_node) + if self.node(conv_node).op_type == "ConvTranspose": + return fold_bn_conv_transpose(conv.weight, conv.bias, bn, groups=conv.group, dtype=dtype) + return fold_bn_conv(conv.weight, conv.bias, bn, dtype=dtype) + + def __repr__(self) -> str: + return f"OnnxWeights({self.path.name}: {len(self._nodes)} nodes, {len(self._tensors)} constants)" + + +# ------------------------------------------------------------------------------------- state dicts + +def _torch_to_numpy(t: Any) -> np.ndarray: + t = t.detach().cpu() + if str(t.dtype) == "torch.bfloat16": + t = t.float() + return t.numpy() + + +def load_safetensors(path: PathLike) -> Dict[str, np.ndarray]: + """All tensors of a ``.safetensors`` file as numpy (bf16 tensors are widened to float32 through torch).""" + from safetensors import safe_open + + out: Dict[str, np.ndarray] = {} + try: + with safe_open(str(path), framework="numpy") as f: + for k in f.keys(): + out[k] = f.get_tensor(k) + return out + except (TypeError, ValueError): # bf16 has no numpy dtype + out.clear() + with safe_open(str(path), framework="pt") as f: + for k in f.keys(): + out[k] = _torch_to_numpy(f.get_tensor(k)) + return out + + +_CHECKPOINT_KEYS = ("state_dict", "model", "model_state_dict", "module", "net", "params") + + +def load_torch_checkpoint(path: PathLike, *, key: Optional[str] = None) -> Dict[str, np.ndarray]: + """A ``.pth`` / ``.pt`` / ``.ckpt`` checkpoint loaded with ``torch.load(weights_only=True)`` (data only: no + pickled code runs) -> ``{name: numpy}``. ``key`` selects a nested dict; otherwise the first of + ``state_dict`` / ``model`` / ``model_state_dict`` / ``module`` / ``net`` / ``params`` that holds tensors.""" + import torch + + obj = torch.load(str(path), map_location="cpu", weights_only=True) + if key is not None: + obj = obj[key] + elif isinstance(obj, Mapping) and not any(hasattr(v, "detach") for v in obj.values()): + for k in _CHECKPOINT_KEYS: + if isinstance(obj.get(k), Mapping): + obj = obj[k] + break + if not isinstance(obj, Mapping): + raise ValueError(f"{path}: no tensor dict found (pass key=)") + return {str(k): _torch_to_numpy(v) for k, v in obj.items() if hasattr(v, "detach")} + + +def load_state_dict(path: PathLike, **kwargs: Any) -> "StateDict": + """``.safetensors`` / ``.pth`` / ``.pt`` / ``.ckpt`` / ``.npz`` -> :class:`StateDict` (by file suffix).""" + p = Path(path) + suffix = p.suffix.lower() + if suffix == ".safetensors": + tensors = load_safetensors(p) + elif suffix in (".pth", ".pt", ".ckpt", ".bin"): + tensors = load_torch_checkpoint(p, **kwargs) + elif suffix == ".npz": + with np.load(p, allow_pickle=False) as z: + tensors = {k: z[k] for k in z.files} + else: + raise ValueError(f"unsupported weights file {p.name}") + return StateDict(tensors) + + +class StateDict(Mapping): + """A read-only ``{name: numpy}`` view with prefix scoping: ``sd.sub("backbone.layer1.")["0.conv1.weight"]``.""" + + def __init__(self, tensors: Mapping[str, np.ndarray], prefix: str = ""): + self._tensors = tensors + self.prefix = prefix + + def __getitem__(self, name: str) -> np.ndarray: + try: + return self._tensors[self.prefix + name] + except KeyError: + raise KeyError(f"{self.prefix + name!r} not in the state dict") from None + + def __iter__(self) -> Iterator[str]: + n = len(self.prefix) + return (k[n:] for k in self._tensors if k.startswith(self.prefix)) + + def __len__(self) -> int: + return sum(1 for _ in self) + + def sub(self, prefix: str) -> "StateDict": + return StateDict(self._tensors, self.prefix + prefix) + + def strip(self, prefix: str) -> "StateDict": + """Drop a leading ``prefix`` (e.g. ``module.``) from every key that has it.""" + return StateDict({(k[len(prefix):] if k.startswith(prefix) else k): v for k, v in self._tensors.items()}, + self.prefix) + + def bn(self, name: str, eps: float = 1e-5) -> BatchNorm: + """BatchNorm ``.{weight,bias,running_mean,running_var}``.""" + return BatchNorm.from_state_dict(self, name, eps) + + +# ----------------------------------------------------------------------------------------- device cache + +def file_sha256(path: PathLike, chunk: int = 1 << 20) -> str: + h = hashlib.sha256() + with open(path, "rb") as f: + for block in iter(lambda: f.read(chunk), b""): + h.update(block) + return h.hexdigest() + + +def cache_root() -> Path: + """``$TT_CACHE_PATH`` (tt-model containers mount ``/tensor-cache`` there), else ``~/.cache/ttaw``.""" + env = os.environ.get("TT_CACHE_PATH") + return Path(env) if env else Path.home() / ".cache" / "ttaw" + + +_SAFE = re.compile(r"[^A-Za-z0-9_.-]+") + + +class WeightCache: + """Prepared device tensors cached on disk as ttnn ``.tensorbin`` files. + + ``namespace`` names the model (``"tt_centerpoint/base"``); ``source_digest`` identifies the source weights + (e.g. ``OnnxWeights.sha256``) so a new revision gets a fresh directory. ``version`` names the preparation + recipe and is part of the directory too: the cache outlives images (tt-model mounts the host's + ``~/.cache/tt-model//tensors`` at ``/tensor-cache``), so an entry built by older preparation code from the + same source weights would otherwise be served to newer code; bump it (e.g. the bundle ``__version__`` plus a + prep revision) whenever the code that builds the tensors changes. Each entry is keyed by name, dtype and layout; + a loaded entry whose dtype, layout or (with ``shape=``) logical shape differs from the request is rebuilt. + ``enabled=None`` reads ``TTAW_WEIGHT_CACHE`` once (``0`` disables: every call rebuilds).""" + + def __init__(self, namespace: str, source_digest: str = "", *, version: str = "", + root: Optional[PathLike] = None, enabled: Optional[bool] = None): + if enabled is None: + enabled = os.environ.get("TTAW_WEIGHT_CACHE", "1").strip() not in ("0", "false", "off", "no") + self.enabled = bool(enabled) + parts = [p for p in namespace.strip("/").split("/") if p] + if not parts or any(p in (".", "..") for p in parts): + raise ValueError(f"bad cache namespace {namespace!r}") + leaf = (source_digest or "nodigest")[:16] + (f"-{_SAFE.sub('_', str(version))}" if version else "") + self.version = str(version) + self.dir = Path(root) if root is not None else cache_root() + self.dir = self.dir.joinpath(*(_SAFE.sub("_", p) for p in parts), leaf) + self.hits = 0 + self.misses = 0 + + def path(self, key: str, dtype: Any, layout: Any) -> Path: + tag = f"{key}|{dtype}|{layout}" + short = hashlib.sha1(tag.encode()).hexdigest()[:8] + dname = getattr(dtype, "name", str(dtype)).lower() + lname = getattr(layout, "name", str(layout)).lower() + return self.dir / f"{_SAFE.sub('_', key)[:120]}.{dname}.{lname}.{short}.tensorbin" + + def get(self, key: str, make: Callable[[], Any], *, dtype: Any, layout: Any = None, device: Any = None, + memory_config: Any = None, shape: Optional[Tuple[int, ...]] = None): + """The cached host tensor for ``key`` (built with ``make()`` -> numpy / torch on a miss), moved to + ``device`` with ``memory_config`` when given. ``shape``: the expected logical shape (an entry of another + shape is rebuilt).""" + import ttnn + + from .tensors import to_host_tensor, ttnn_dtype + + dtype = ttnn_dtype(dtype) + layout = ttnn.TILE_LAYOUT if layout is None else layout + host = None + p = self.path(key, dtype, layout) + if self.enabled and p.is_file(): + try: + host = ttnn.load_tensor(p) + except Exception: # noqa: BLE001 -- unreadable / old format: rebuild below + host = None + if host is not None and (host.dtype != dtype or host.layout != layout + or (shape is not None and tuple(host.shape) != tuple(shape))): + host = None # an entry of another recipe: rebuild + if host is not None: + self.hits += 1 + if host is None: + self.misses += 1 + host = to_host_tensor(make(), dtype, layout, shape=shape) + if self.enabled: + p.parent.mkdir(parents=True, exist_ok=True) + tmp = p.with_name(f".{p.stem}.{os.getpid()}.tmp.tensorbin") + ttnn.dump_tensor(tmp, host) + os.replace(tmp, p) + if device is None: + return host + return ttnn.to_device(host, device, memory_config=memory_config or ttnn.DRAM_MEMORY_CONFIG) + + def clear(self) -> int: + """Delete this namespace/digest directory's entries; returns how many files were removed.""" + n = 0 + if self.dir.is_dir(): + for f in self.dir.glob("*.tensorbin"): + f.unlink() + n += 1 + return n diff --git a/examples/quickstart.py b/examples/quickstart.py new file mode 100644 index 0000000000000000000000000000000000000000..2e4e862d281e501cb9f8d477b20a0f87d7636d33 --- /dev/null +++ b/examples/quickstart.py @@ -0,0 +1,100 @@ +#!/usr/bin/env python3 +# SPDX-License-Identifier: Apache-2.0 +"""Quickstart: the model card's Python snippet on the shipped sample, on one Blackhole p150. + + pip install -e . # once, from the repo root, on top of an environment that has ttnn (tt-metal) + python examples/quickstart.py [input.npz] [--out-dir examples/output] + +Writes /quickstart.json (the same JSON as POST /predict: the 8 s ego trajectory, the turn-indicator command +and the predicted paths of the neighbours) and /quickstart_bev.png (a bird's-eye view of the input tensors +with the plan: lanes, route lanes, stop lines and road borders, the neighbours with their predicted 8 s paths, the ego +plan with a dot every second). The default input is found relative to this file (runs from any directory); an input +given on the command line is relative to the current directory and must hold the 15 raw planner tensors +(`DiffusionPlanner.INPUT_SCHEMA`). +""" +import argparse +import json +from pathlib import Path + +REPO = Path(__file__).resolve().parents[1] +ap = argparse.ArgumentParser() +ap.add_argument("input", nargs="?", + default=str(REPO / "code" / "tt_diffusion_planner" / "samples" / "kashiwanoha_dense.npz")) +ap.add_argument("--out-dir", default=str(REPO / "examples" / "output")) +ap.add_argument("--device-id", type=int, default=0) +args = ap.parse_args() +out_dir = Path(args.out_dir) +out_dir.mkdir(parents=True, exist_ok=True) + +# --- the model card snippet -------------------------------------------------------------------------------------- +from tt_diffusion_planner import DiffusionPlanner + +with DiffusionPlanner.from_pretrained(device_id=args.device_id) as model: # weights -> your HF cache, traces captured + out = model(inputs=args.input) # the 15 raw planner tensors: .npz path, its bytes, or {name: array} + +print(out.columns) # x, y, yaw, cos, sin, velocity, acceleration (base_link, 0.1-8.0 s) +print(out.poses[:5]) +print(out.turn_indicator["command_name"], out.predicted_agents.shape) +# ------------------------------------------------------------------------------------------------------------------ + +(out_dir / "quickstart.json").write_text(json.dumps(out.to_dict(), indent=1)) + +# A bird's-eye view of what the model saw and what it planned (ego frame: x forward, y left; pillow only). +import numpy as np # noqa: E402 +from PIL import Image, ImageDraw # noqa: E402 + +from tt_diffusion_planner import load_inputs # noqa: E402 + +raw = {k: v[0] for k, v in load_inputs(args.input).items()} # the same decoder as model(inputs=...) +plan = out.poses[:, :2] +fwd = max(60.0, float(plan[:, 0].max()) + 15.0) +x0, x1 = -20.0, fwd # forward range (m) +half = (x1 - x0) / 2 # lateral half-width (m): a square view +W = H = 800 +s = W / (2 * half) # pixels per metre + + +def px(xy): + """ego-frame metres [N, 2] -> image pixels (forward up, left to the left).""" + xy = np.asarray(xy, np.float64).reshape(-1, 2) + return [(W / 2 - y * s, H - (x - x0) * s) for x, y in xy] + + +img = Image.new("RGB", (W, H), (252, 252, 251)) +d = ImageDraw.Draw(img) +for name, fill in (("lanes", (236, 235, 231)), ("route_lanes", (205, 226, 251))): + t = raw[name] + for lane in t[np.abs(t[:, :, :8]).sum(axis=(1, 2)) > 0]: + left, right = lane[:, :2] + lane[:, 4:6], lane[:, :2] + lane[:, 6:8] + d.polygon(px(np.concatenate([left, right[::-1]])), fill=fill) +for lane in raw["lanes"][np.abs(raw["lanes"][:, :, :8]).sum(axis=(1, 2)) > 0]: + for off in (4, 6): + d.line(px(lane[:, :2] + lane[:, off:off + 2]), fill=(195, 194, 183), width=1) +for ls in raw["line_strings"][np.abs(raw["line_strings"]).sum(axis=(1, 2)) > 0]: + stop = ls[0, 2] > 0.5 # line-string type: stop line, else road border + d.line(px(ls[:, :2]), fill=(11, 11, 11) if stop else (82, 81, 78), width=4 if stop else 2) +colors = {8: (235, 104, 52), 9: (232, 123, 164), 10: (27, 175, 122)} # vehicle, pedestrian, bicycle +nb = raw["neighbor_agents_past"] +paths = dict(zip(out.meta["predicted_agent_rows"], out.predicted_agents)) +for i in np.flatnonzero(np.abs(nb[:, -1, :8]).sum(axis=1) > 0): + x, y, c, sn, w, length = (float(v) for v in nb[i, -1, [0, 1, 2, 3, 6, 7]]) + col = colors[8 + int(np.argmax(nb[i, -1, 8:11]))] + if i in paths: + d.line(px(np.concatenate([[[x, y]], paths[i][:, :2]])), fill=col, width=1) + f, lt = np.array([c, sn]), np.array([-sn, c]) + corners = [np.array([x, y]) + a * max(length, 0.5) / 2 * f + b * max(w, 0.5) / 2 * lt + for a, b in ((1, 1), (1, -1), (-1, -1), (-1, 1))] + d.polygon(px(corners), fill=col) +wb, length, w = (float(v) for v in raw["ego_shape"]) # wheel base, length, width; base_link = rear axle centre +r = (length - wb) / 2 +d.polygon(px([(wb + r, w / 2), (wb + r, -w / 2), (-r, -w / 2), (-r, w / 2)]), fill=(11, 11, 11)) # ego +d.line(px(np.concatenate([[[0.0, 0.0]], plan])), fill=(42, 120, 214), width=4) +for u, v in px(plan[9::10]): # one dot per second + d.ellipse([u - 5, v - 5, u + 5, v + 5], fill=(42, 120, 214), outline=(252, 252, 251), width=2) +d.text((10, 8), f"{Path(args.input).name}: 8 s ego plan (blue, a dot per second), " + f"turn indicator {out.turn_indicator['command_name']}", fill=(11, 11, 11)) +d.text((10, 24), f"{len(paths)} neighbours with predicted 8 s paths; view {2 * half:.0f} m wide, ego frame " + "(forward up)", fill=(82, 81, 78)) +img.save(out_dir / "quickstart_bev.png") +print(f"{len(out)} poses, turn {out.turn_indicator['command_name']} -> {out_dir / 'quickstart.json'}, " + f"{out_dir / 'quickstart_bev.png'} timing_ms={ {k: round(v, 2) for k, v in out.timing_ms.items()} }") diff --git a/image/blobs/sha256/069395655fd57335d9426528f7f49731be6cf7ed963ca649e4ff7c9697d37e0a b/image/blobs/sha256/069395655fd57335d9426528f7f49731be6cf7ed963ca649e4ff7c9697d37e0a new file mode 100644 index 0000000000000000000000000000000000000000..8fcbc96dafe31cbf7242c8f63266528012400ada --- /dev/null +++ b/image/blobs/sha256/069395655fd57335d9426528f7f49731be6cf7ed963ca649e4ff7c9697d37e0a @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:069395655fd57335d9426528f7f49731be6cf7ed963ca649e4ff7c9697d37e0a +size 63869034 diff --git a/image/blobs/sha256/0936b6c31255cfe6d029dab00fc251d8771f5e3d228ac7da23e395e1161c6992 b/image/blobs/sha256/0936b6c31255cfe6d029dab00fc251d8771f5e3d228ac7da23e395e1161c6992 new file mode 100644 index 0000000000000000000000000000000000000000..f2b0607bcbdf190a0591bca5773d9a7874e5edd0 --- /dev/null +++ b/image/blobs/sha256/0936b6c31255cfe6d029dab00fc251d8771f5e3d228ac7da23e395e1161c6992 @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:0936b6c31255cfe6d029dab00fc251d8771f5e3d228ac7da23e395e1161c6992 +size 41057719 diff --git a/image/blobs/sha256/18c4be700a3778a5b45eec8a3647e95f4376c403394168df7f0251356baad738 b/image/blobs/sha256/18c4be700a3778a5b45eec8a3647e95f4376c403394168df7f0251356baad738 new file mode 100644 index 0000000000000000000000000000000000000000..e85f5ad0f83a65b8cd24e9cacb3e10a0a63c0251 Binary files /dev/null and b/image/blobs/sha256/18c4be700a3778a5b45eec8a3647e95f4376c403394168df7f0251356baad738 differ diff --git a/image/blobs/sha256/3ac9de4540f40bdf99f054641b2fb36a38f49253d1727fb08a73272d4208ce1b b/image/blobs/sha256/3ac9de4540f40bdf99f054641b2fb36a38f49253d1727fb08a73272d4208ce1b new file mode 100644 index 0000000000000000000000000000000000000000..7de558efca421a540b07711cd6fcd3d2ac23802e --- /dev/null +++ b/image/blobs/sha256/3ac9de4540f40bdf99f054641b2fb36a38f49253d1727fb08a73272d4208ce1b @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:3ac9de4540f40bdf99f054641b2fb36a38f49253d1727fb08a73272d4208ce1b +size 66024983 diff --git a/image/blobs/sha256/3b96d8ea71902fe6a00f1792dd41290839b7758070431464cd613cde3f6bf909 b/image/blobs/sha256/3b96d8ea71902fe6a00f1792dd41290839b7758070431464cd613cde3f6bf909 new file mode 100644 index 0000000000000000000000000000000000000000..fc0fe7abdd8d5b76e5919dc8732efc54d64f79a8 --- /dev/null +++ b/image/blobs/sha256/3b96d8ea71902fe6a00f1792dd41290839b7758070431464cd613cde3f6bf909 @@ -0,0 +1,28 @@ +{ + "schemaVersion": 2, + "mediaType": "application/vnd.oci.image.index.v1+json", + "manifests": [ + { + "mediaType": "application/vnd.oci.image.manifest.v1+json", + "digest": "sha256:fd2cd59f930cf2becfb547624bad96b2391338e765be2c76d5e16660a77ac598", + "size": 4684, + "platform": { + "architecture": "amd64", + "os": "linux" + } + }, + { + "mediaType": "application/vnd.oci.image.manifest.v1+json", + "digest": "sha256:cb04bedc6712d6c313997412aafab6475bde03c65fce33faf5918691a6978fc4", + "size": 837, + "annotations": { + "vnd.docker.reference.digest": "sha256:fd2cd59f930cf2becfb547624bad96b2391338e765be2c76d5e16660a77ac598", + "vnd.docker.reference.type": "attestation-manifest" + }, + "platform": { + "architecture": "unknown", + "os": "unknown" + } + } + ] +} \ No newline at end of file diff --git a/image/blobs/sha256/44136fa355b3678a1146ad16f7e8649e94fb4fc21fe77e8310c060f61caaff8a b/image/blobs/sha256/44136fa355b3678a1146ad16f7e8649e94fb4fc21fe77e8310c060f61caaff8a new file mode 100644 index 0000000000000000000000000000000000000000..9e26dfeeb6e641a33dae4961196235bdb965b21b --- /dev/null +++ b/image/blobs/sha256/44136fa355b3678a1146ad16f7e8649e94fb4fc21fe77e8310c060f61caaff8a @@ -0,0 +1 @@ +{} \ No newline at end of file diff --git a/image/blobs/sha256/46f21edbef18d35cb9e2f400eeb6d8f8fac5873f8620e8d96829a9c01447ed52 b/image/blobs/sha256/46f21edbef18d35cb9e2f400eeb6d8f8fac5873f8620e8d96829a9c01447ed52 new file mode 100644 index 0000000000000000000000000000000000000000..533174d3c1426726d8671b4136aaa6e407f6e966 --- /dev/null +++ b/image/blobs/sha256/46f21edbef18d35cb9e2f400eeb6d8f8fac5873f8620e8d96829a9c01447ed52 @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:46f21edbef18d35cb9e2f400eeb6d8f8fac5873f8620e8d96829a9c01447ed52 +size 139083872 diff --git a/image/blobs/sha256/4ad0ad64dc956c94bdcf581b56664b50f1d1ac6d091e6a7b1f28bf612bd7ab69 b/image/blobs/sha256/4ad0ad64dc956c94bdcf581b56664b50f1d1ac6d091e6a7b1f28bf612bd7ab69 new file mode 100644 index 0000000000000000000000000000000000000000..2d071b381396b3c598428ef151985774d1e2e83a Binary files /dev/null and b/image/blobs/sha256/4ad0ad64dc956c94bdcf581b56664b50f1d1ac6d091e6a7b1f28bf612bd7ab69 differ diff --git a/image/blobs/sha256/4f4fb700ef54461cfa02571ae0db9a0dc1e0cdb5577484a6d75e68dc38e8acc1 b/image/blobs/sha256/4f4fb700ef54461cfa02571ae0db9a0dc1e0cdb5577484a6d75e68dc38e8acc1 new file mode 100644 index 0000000000000000000000000000000000000000..8de868223f695a02a4da72b1e79a6545868048c5 Binary files /dev/null and b/image/blobs/sha256/4f4fb700ef54461cfa02571ae0db9a0dc1e0cdb5577484a6d75e68dc38e8acc1 differ diff --git a/image/blobs/sha256/8bc4b6829641a7f2dde4ebe50b8189de860e7ec898b8570f0922a540b3c0ca6a b/image/blobs/sha256/8bc4b6829641a7f2dde4ebe50b8189de860e7ec898b8570f0922a540b3c0ca6a new file mode 100644 index 0000000000000000000000000000000000000000..83263db012942fde53245eca7cb9fd25ef1e088a --- /dev/null +++ b/image/blobs/sha256/8bc4b6829641a7f2dde4ebe50b8189de860e7ec898b8570f0922a540b3c0ca6a @@ -0,0 +1 @@ +{"architecture":"amd64","config":{"User":"tt","Env":["PATH=/opt/tt-venv/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin","VENV=/opt/tt-venv","VIRTUAL_ENV=/opt/tt-venv","TT_METAL_RUNTIME_ROOT=/opt/tt-metal","TT_METAL_HOME=/opt/tt-metal","PYTHONPATH=/opt/tt-metal","LD_LIBRARY_PATH=/opt/tt-metal/build/lib:/opt/openmpi-v5.0.7-ulfm/lib","EXTRA_MODELS_DIR=","TT_VLLM_BUILTIN_MODELS=","TT_MODEL_KIND=tt-dit-server","HF_HOME=/hf","TT_METAL_CACHE=/cache","HOME=/home/tt","USER=tt","LOGNAME=tt"],"Entrypoint":["/usr/local/bin/entrypoint.sh"],"Cmd":["/usr/local/bin/serve-default.sh"],"WorkingDir":"/home/tt/work","Labels":{"org.opencontainers.image.revision":"44d66500520fda9f2c7060c0f6b41ec48f7ab37e","org.opencontainers.image.version":"22.04","org.tenstorrent.tt-model":"diffusion-planner-p150","org.tenstorrent.tt-model.arch":"blackhole","org.tenstorrent.tt-model.kind":"tt-dit-server","org.tenstorrent.tt-model.plugin":"","org.tenstorrent.tt-model.profiles":"default","org.tenstorrent.tt-model.repo":"changh95/diffusion-planner-p150","org.tenstorrent.tt-model.tt-metal":"v0.80.0-dev20261006-78-g44d6650052-dirty","org.tenstorrent.tt-model.weights":"AutowareFoundation/diffusion_planner"},"ArgsEscaped":true},"created":"2026-10-09T04:46:06.252634349Z","history":[{"created":"2026-09-24T22:24:04.688908979Z","created_by":"/bin/sh -c #(nop) ARG RELEASE","empty_layer":true},{"created":"2026-09-24T22:24:04.723527376Z","created_by":"/bin/sh -c #(nop) ARG LAUNCHPAD_BUILD_ARCH","empty_layer":true},{"created":"2026-09-24T22:24:04.751455651Z","created_by":"/bin/sh -c #(nop) LABEL org.opencontainers.image.version=22.04","empty_layer":true},{"created":"2026-09-24T22:24:06.837737736Z","created_by":"/bin/sh -c #(nop) ADD file:b0bf3f64519bf10a51e00d4f8ab9c8693620659a0c4a4a5090a9765e754e1d1c in / "},{"created":"2026-09-24T22:24:07.239071539Z","created_by":"/bin/sh -c #(nop) CMD [\"/bin/bash\"]","empty_layer":true},{"created":"2026-10-09T04:37:45.428325056Z","created_by":"ARG OMPI_DIR=/opt/openmpi-v5.0.7-ulfm","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:37:45.428325056Z","created_by":"ARG EXTRA_MODELS_DIR=","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:37:45.428325056Z","created_by":"ARG TT_MODEL_KIND","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:37:45.428325056Z","created_by":"ARG MODEL_NAME","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:37:45.428325056Z","created_by":"ARG MODEL_REPO","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:37:45.428325056Z","created_by":"ARG MODEL_WEIGHTS","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:37:45.428325056Z","created_by":"ARG MODEL_ARCH","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:37:45.428325056Z","created_by":"ARG MODEL_PROFILES","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:37:45.428325056Z","created_by":"ARG MODEL_TT_METAL_SHA","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:37:45.428325056Z","created_by":"ARG MODEL_TT_METAL_DESCRIBE","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:37:45.428325056Z","created_by":"ARG MODEL_PLUGIN_SHA","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:37:45.428325056Z","created_by":"RUN |11 OMPI_DIR=/opt/openmpi-v5.0.7-ulfm EXTRA_MODELS_DIR= TT_MODEL_KIND=tt-dit-server MODEL_NAME=diffusion-planner-p150 MODEL_REPO=changh95/diffusion-planner-p150 MODEL_WEIGHTS=AutowareFoundation/diffusion_planner MODEL_ARCH=blackhole MODEL_PROFILES=default MODEL_TT_METAL_SHA=44d66500520fda9f2c7060c0f6b41ec48f7ab37e MODEL_TT_METAL_DESCRIBE=v0.80.0-dev20261006-78-g44d6650052-dirty MODEL_PLUGIN_SHA= /bin/sh -c apt-get update \u0026\u0026 apt-get install -y --no-install-recommends libhwloc15 libnuma1 libatomic1 libudev1 libcap2 zlib1g libmpc3 libmpfr6 libgmp10 libzstd1 libevent-core-2.1-7 libevent-pthreads-2.1-7 libgl1 libsndfile1 ca-certificates \u0026\u0026 apt-get clean \u0026\u0026 rm -rf /var/lib/apt/lists/* # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-09T04:37:47.354882893Z","created_by":"RUN |11 OMPI_DIR=/opt/openmpi-v5.0.7-ulfm EXTRA_MODELS_DIR= TT_MODEL_KIND=tt-dit-server MODEL_NAME=diffusion-planner-p150 MODEL_REPO=changh95/diffusion-planner-p150 MODEL_WEIGHTS=AutowareFoundation/diffusion_planner MODEL_ARCH=blackhole MODEL_PROFILES=default MODEL_TT_METAL_SHA=44d66500520fda9f2c7060c0f6b41ec48f7ab37e MODEL_TT_METAL_DESCRIBE=v0.80.0-dev20261006-78-g44d6650052-dirty MODEL_PLUGIN_SHA= /bin/sh -c existing=\"$(getent passwd 1000 | cut -d: -f1)\" \u0026\u0026 if [ -n \"$existing\" ]; then userdel -r \"$existing\" 2\u003e/dev/null || userdel \"$existing\"; fi \u0026\u0026 useradd --uid 1000 --create-home --home-dir /home/tt --shell /bin/bash tt \u0026\u0026 mkdir -p /home/tt/work/logs /cache /opt/tt-metal \u0026\u0026 chown -R tt:tt /home/tt /cache /opt/tt-metal \u0026\u0026 chmod 1777 /home/tt /home/tt/work /home/tt/work/logs /cache # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-09T04:45:25.944507344Z","created_by":"COPY /opt/openmpi-v5.0.7-ulfm /opt/openmpi-v5.0.7-ulfm # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-09T04:45:27.594701615Z","created_by":"COPY /opt/tenstorrent /opt/tenstorrent # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-09T04:45:30.393973408Z","created_by":"COPY /usr/local/share/uv /usr/local/share/uv # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-09T04:45:37.451313045Z","created_by":"COPY /opt/tt-venv /opt/tt-venv # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-09T04:45:37.701138787Z","created_by":"COPY /opt/vllm /opt/vllm # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-09T04:45:39.148213219Z","created_by":"COPY --chown=tt:tt /opt/tt-metal/runtime /opt/tt-metal/runtime # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-09T04:45:39.577641501Z","created_by":"COPY --chown=tt:tt /opt/tt-metal/build_Release /opt/tt-metal/build_Release # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-09T04:45:39.991364785Z","created_by":"COPY --chown=tt:tt /opt/tt-metal/build /opt/tt-metal/build # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-09T04:45:41.810332453Z","created_by":"COPY --chown=tt:tt /opt/tt-metal/tt_metal /opt/tt-metal/tt_metal # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-09T04:45:43.353759095Z","created_by":"COPY --chown=tt:tt /opt/tt-metal/ttnn /opt/tt-metal/ttnn # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-09T04:45:43.568335993Z","created_by":"COPY --chown=tt:tt /opt/tt-metal/tools /opt/tt-metal/tools # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-09T04:45:43.610659256Z","created_by":"COPY --chown=tt:tt /opt/tt-metal/setup.py /opt/tt-metal/pyproject.toml /opt/tt-metal/ # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-09T04:45:43.663189757Z","created_by":"COPY --chown=tt:tt code/ /opt/tt-metal/ # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-09T04:45:44.209653476Z","created_by":"RUN |11 OMPI_DIR=/opt/openmpi-v5.0.7-ulfm EXTRA_MODELS_DIR= TT_MODEL_KIND=tt-dit-server MODEL_NAME=diffusion-planner-p150 MODEL_REPO=changh95/diffusion-planner-p150 MODEL_WEIGHTS=AutowareFoundation/diffusion_planner MODEL_ARCH=blackhole MODEL_PROFILES=default MODEL_TT_METAL_SHA=44d66500520fda9f2c7060c0f6b41ec48f7ab37e MODEL_TT_METAL_DESCRIBE=v0.80.0-dev20261006-78-g44d6650052-dirty MODEL_PLUGIN_SHA= /bin/sh -c find /opt/tt-metal \\( -type f -o -type d \\) \\( ! -perm -o+r -o \\( -type d ! -perm -o+x \\) \\) -exec chmod a+rX {} + # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-09T04:45:44.237742415Z","created_by":"COPY --chmod=0755 entrypoint.sh /usr/local/bin/entrypoint.sh # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-09T04:45:44.260006284Z","created_by":"COPY --chmod=0755 serve-default.sh /usr/local/bin/serve-default.sh # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-09T04:45:44.260006284Z","created_by":"ENV VENV=/opt/tt-venv","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:45:44.260006284Z","created_by":"ENV VIRTUAL_ENV=/opt/tt-venv","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:45:44.260006284Z","created_by":"ENV PATH=/opt/tt-venv/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:45:44.260006284Z","created_by":"ENV TT_METAL_RUNTIME_ROOT=/opt/tt-metal","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:45:44.260006284Z","created_by":"ENV TT_METAL_HOME=/opt/tt-metal","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:45:44.260006284Z","created_by":"ENV PYTHONPATH=/opt/tt-metal","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:45:44.260006284Z","created_by":"ENV LD_LIBRARY_PATH=/opt/tt-metal/build/lib:/opt/openmpi-v5.0.7-ulfm/lib","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:45:44.260006284Z","created_by":"ENV EXTRA_MODELS_DIR=","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:45:44.260006284Z","created_by":"ARG TT_VLLM_BUILTIN_MODELS=","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:45:44.260006284Z","created_by":"ENV TT_VLLM_BUILTIN_MODELS=","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:45:44.260006284Z","created_by":"ENV TT_MODEL_KIND=tt-dit-server","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:45:44.260006284Z","created_by":"ENV HF_HOME=/hf","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:45:44.260006284Z","created_by":"ENV TT_METAL_CACHE=/cache","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:45:44.260006284Z","created_by":"ENV HOME=/home/tt","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:45:44.260006284Z","created_by":"ENV USER=tt","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:45:44.260006284Z","created_by":"ENV LOGNAME=tt","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:45:44.260006284Z","created_by":"USER tt","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:45:44.284612834Z","created_by":"WORKDIR /home/tt/work","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-09T04:45:44.306849883Z","created_by":"COPY --chmod=0755 verify.sh /ctx/verify.sh # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-09T04:46:06.033742536Z","created_by":"RUN |12 OMPI_DIR=/opt/openmpi-v5.0.7-ulfm EXTRA_MODELS_DIR= TT_MODEL_KIND=tt-dit-server MODEL_NAME=diffusion-planner-p150 MODEL_REPO=changh95/diffusion-planner-p150 MODEL_WEIGHTS=AutowareFoundation/diffusion_planner MODEL_ARCH=blackhole MODEL_PROFILES=default MODEL_TT_METAL_SHA=44d66500520fda9f2c7060c0f6b41ec48f7ab37e MODEL_TT_METAL_DESCRIBE=v0.80.0-dev20261006-78-g44d6650052-dirty MODEL_PLUGIN_SHA= TT_VLLM_BUILTIN_MODELS= /bin/sh -c bash /ctx/verify.sh # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-09T04:46:06.033742536Z","created_by":"USER root","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:46:06.252634349Z","created_by":"RUN |12 OMPI_DIR=/opt/openmpi-v5.0.7-ulfm EXTRA_MODELS_DIR= TT_MODEL_KIND=tt-dit-server MODEL_NAME=diffusion-planner-p150 MODEL_REPO=changh95/diffusion-planner-p150 MODEL_WEIGHTS=AutowareFoundation/diffusion_planner MODEL_ARCH=blackhole MODEL_PROFILES=default MODEL_TT_METAL_SHA=44d66500520fda9f2c7060c0f6b41ec48f7ab37e MODEL_TT_METAL_DESCRIBE=v0.80.0-dev20261006-78-g44d6650052-dirty MODEL_PLUGIN_SHA= TT_VLLM_BUILTIN_MODELS= /bin/sh -c chmod -R a+rwX /home/tt # buildkit","comment":"buildkit.dockerfile.v0"},{"created":"2026-10-09T04:46:06.252634349Z","created_by":"USER tt","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:46:06.252634349Z","created_by":"LABEL org.tenstorrent.tt-model=diffusion-planner-p150 org.tenstorrent.tt-model.repo=changh95/diffusion-planner-p150 org.tenstorrent.tt-model.weights=AutowareFoundation/diffusion_planner org.tenstorrent.tt-model.arch=blackhole org.tenstorrent.tt-model.kind=tt-dit-server org.tenstorrent.tt-model.profiles=default org.opencontainers.image.revision=44d66500520fda9f2c7060c0f6b41ec48f7ab37e org.tenstorrent.tt-model.tt-metal=v0.80.0-dev20261006-78-g44d6650052-dirty org.tenstorrent.tt-model.plugin=","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:46:06.252634349Z","created_by":"ENTRYPOINT [\"/usr/local/bin/entrypoint.sh\"]","comment":"buildkit.dockerfile.v0","empty_layer":true},{"created":"2026-10-09T04:46:06.252634349Z","created_by":"CMD [\"/usr/local/bin/serve-default.sh\"]","comment":"buildkit.dockerfile.v0","empty_layer":true}],"os":"linux","rootfs":{"type":"layers","diff_ids":["sha256:cbaf391e670933f026c05c7dedec3faf8f194b6c2d83b0d5694ad68323c8077c","sha256:a37a1ce75193138998bd1453ef307cb02b0ccc63031cde4037e9a173912ffb9b","sha256:ae9c33bdf8e967407b896117a7df1e16033aaa1e111a88feda14bf00088d45a9","sha256:a491a9bc0ecd491ca7e8e6e21f1b26967c3fc0e1d6444b2fd2fcece0c4332e5d","sha256:34d4b96347ac8a9feb4e324c34394eea1aee623e829aee4dd0506788e9597094","sha256:09e9040caf7b6a2721e83eef2f6e445c71603daa2d5475e77aef9af9db579d10","sha256:1ddba2587aae8e309603ba767a596b140156b87de6d043562eca7bd63ad17e93","sha256:68b69108241e57e85e31f6ebf4b0a06358c566b50bc868850fa44511de9c8db0","sha256:08bb9ef30137c0a766982267db392c721e585db43c0cd2cf2cca3e2cd6aa7198","sha256:49370605aea38c1ef103b99a25c1c7f4c92661a29602a6d73d0617c7aa252003","sha256:05e164bf39f278254ca5c77723b31ab8d151be48d336b66a4f1565509bc1e68b","sha256:6efbe280729d314f8edfeca0c29c47238f8e1b2c42ca88da994701a50b3cc3db","sha256:81236f19b07e8bb5c9399a6149bbe7b89390b9adf3d09bfe3e1bd485cadf8038","sha256:6a178e1f45ce4394add0f0e49dcc315f074aa451f6831d90e79fef842ab3fe05","sha256:bc43cd5c0e960c964d52072ce492c6d0ecdcba8ff5e00e18b8ca795e2c13c53e","sha256:0d1cd6283ddf6f7784f9fbf2aef38af0037421465befa74062b0e7c57e96b745","sha256:5f70bf18a086007016e948b04aed3b82103a36bea41755b6cddfaf10ace3c6ef","sha256:7c4b12554b6b0a59143ac41b817cc0fd68c1679d2b81322b70638ee4c1492cc3","sha256:b933076c5576abdb652f62caac2f327c3347fa84bb1cdd64a30ff89c0482a9f4","sha256:5f70bf18a086007016e948b04aed3b82103a36bea41755b6cddfaf10ace3c6ef","sha256:d2387c89efa220d8d121f73353aad0dea53578bcc1de05658e3fbbfa2a34fd10","sha256:ba9922c2bbc08e9dcba9460920415f5437a4404ddf3a4d4194d631655bf92c25","sha256:ffca2d722e8b95d1b782c1b95e6a7fbe987aa57943f91eeee881af72231d8cde"]}} \ No newline at end of file diff --git a/image/blobs/sha256/98c4455a98982b35380ec62f90d5cf0edff02eeb4b3bb18595d871fdeb0c0c67 b/image/blobs/sha256/98c4455a98982b35380ec62f90d5cf0edff02eeb4b3bb18595d871fdeb0c0c67 new file mode 100644 index 0000000000000000000000000000000000000000..e6398658b5944b13a89f9749941a4db58c6f9d46 --- /dev/null +++ b/image/blobs/sha256/98c4455a98982b35380ec62f90d5cf0edff02eeb4b3bb18595d871fdeb0c0c67 @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:98c4455a98982b35380ec62f90d5cf0edff02eeb4b3bb18595d871fdeb0c0c67 +size 29751627 diff --git a/image/blobs/sha256/a15921f6aeba449853d8dfc7215c466ad00cff30329183f8b839522d33bc6552 b/image/blobs/sha256/a15921f6aeba449853d8dfc7215c466ad00cff30329183f8b839522d33bc6552 new file mode 100644 index 0000000000000000000000000000000000000000..786bccd846e143df4f9e0fc2b0c928f6624dbad2 --- /dev/null +++ b/image/blobs/sha256/a15921f6aeba449853d8dfc7215c466ad00cff30329183f8b839522d33bc6552 @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:a15921f6aeba449853d8dfc7215c466ad00cff30329183f8b839522d33bc6552 +size 9565630 diff --git a/image/blobs/sha256/a776bb9879e89222f1941fe93976727dbbfd5e8f5da8f65f42b05095ecb25808 b/image/blobs/sha256/a776bb9879e89222f1941fe93976727dbbfd5e8f5da8f65f42b05095ecb25808 new file mode 100644 index 0000000000000000000000000000000000000000..189ba62e8b2498ee7d9c53a9f35e8c16b2c711ab Binary files /dev/null and b/image/blobs/sha256/a776bb9879e89222f1941fe93976727dbbfd5e8f5da8f65f42b05095ecb25808 differ diff --git a/image/blobs/sha256/a897c24f6585a41459698990eabb275e716204a7310044da5cfda2b0f61809ee b/image/blobs/sha256/a897c24f6585a41459698990eabb275e716204a7310044da5cfda2b0f61809ee new file mode 100644 index 0000000000000000000000000000000000000000..f58d8b56ebdb5f41db1be645b27568ddbfe2922d Binary files /dev/null and b/image/blobs/sha256/a897c24f6585a41459698990eabb275e716204a7310044da5cfda2b0f61809ee differ diff --git a/image/blobs/sha256/ab2ce486a4537408cb227812400260a4fe3e3c19e860f0a2c65aae2a8d889f86 b/image/blobs/sha256/ab2ce486a4537408cb227812400260a4fe3e3c19e860f0a2c65aae2a8d889f86 new file mode 100644 index 0000000000000000000000000000000000000000..50886d0c5a35845201325ee43b161b7c3c2d6b70 Binary files /dev/null and b/image/blobs/sha256/ab2ce486a4537408cb227812400260a4fe3e3c19e860f0a2c65aae2a8d889f86 differ diff --git a/image/blobs/sha256/be04a2af1cfe08908467229c0dd35fc629cb65ba98a497e7b627499f0c76821b b/image/blobs/sha256/be04a2af1cfe08908467229c0dd35fc629cb65ba98a497e7b627499f0c76821b new file mode 100644 index 0000000000000000000000000000000000000000..9deea75ab3ec1be9b9c7bbec474548c83ec02c1a --- /dev/null +++ b/image/blobs/sha256/be04a2af1cfe08908467229c0dd35fc629cb65ba98a497e7b627499f0c76821b @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:be04a2af1cfe08908467229c0dd35fc629cb65ba98a497e7b627499f0c76821b +size 138968702 diff --git a/image/blobs/sha256/c006f8f8bd1ac6c419a231b5d94521d87d03b141b76873d7a4d7fb68a933e57a b/image/blobs/sha256/c006f8f8bd1ac6c419a231b5d94521d87d03b141b76873d7a4d7fb68a933e57a new file mode 100644 index 0000000000000000000000000000000000000000..f87a4ea5e4b4471e2aad87f94f7f3d7512dc632e --- /dev/null +++ b/image/blobs/sha256/c006f8f8bd1ac6c419a231b5d94521d87d03b141b76873d7a4d7fb68a933e57a @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:c006f8f8bd1ac6c419a231b5d94521d87d03b141b76873d7a4d7fb68a933e57a +size 848352 diff --git a/image/blobs/sha256/c74a07e1fb0aaf18cdccf3f41385a21dd137109c155363830e2ac2c565d7ac43 b/image/blobs/sha256/c74a07e1fb0aaf18cdccf3f41385a21dd137109c155363830e2ac2c565d7ac43 new file mode 100644 index 0000000000000000000000000000000000000000..a66eb57cb41976172bfcfb751d5c35422bc8b3cd --- /dev/null +++ b/image/blobs/sha256/c74a07e1fb0aaf18cdccf3f41385a21dd137109c155363830e2ac2c565d7ac43 @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:c74a07e1fb0aaf18cdccf3f41385a21dd137109c155363830e2ac2c565d7ac43 +size 1949486 diff --git a/image/blobs/sha256/ca0cef561ac457e6f5fe6e064c167bfc523b6e975437121969a8f39630a0a424 b/image/blobs/sha256/ca0cef561ac457e6f5fe6e064c167bfc523b6e975437121969a8f39630a0a424 new file mode 100644 index 0000000000000000000000000000000000000000..e8819feedb7145d7a6a205e229b4992d8a826f77 --- /dev/null +++ b/image/blobs/sha256/ca0cef561ac457e6f5fe6e064c167bfc523b6e975437121969a8f39630a0a424 @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:ca0cef561ac457e6f5fe6e064c167bfc523b6e975437121969a8f39630a0a424 +size 41057991 diff --git a/image/blobs/sha256/cb04bedc6712d6c313997412aafab6475bde03c65fce33faf5918691a6978fc4 b/image/blobs/sha256/cb04bedc6712d6c313997412aafab6475bde03c65fce33faf5918691a6978fc4 new file mode 100644 index 0000000000000000000000000000000000000000..1c529145e8e6f531b5d4b0b75774d11d67c1e666 --- /dev/null +++ b/image/blobs/sha256/cb04bedc6712d6c313997412aafab6475bde03c65fce33faf5918691a6978fc4 @@ -0,0 +1,26 @@ +{ + "schemaVersion": 2, + "mediaType": "application/vnd.oci.image.manifest.v1+json", + "artifactType": "application/vnd.docker.attestation.manifest.v1+json", + "config": { + "mediaType": "application/vnd.oci.empty.v1+json", + "digest": "sha256:44136fa355b3678a1146ad16f7e8649e94fb4fc21fe77e8310c060f61caaff8a", + "size": 2, + "data": "e30=" + }, + "layers": [ + { + "mediaType": "application/vnd.in-toto+json", + "digest": "sha256:ed9f975b7148483c034ca6dec37793ef0d6ed4d60b99159a3dcd44fd56c847fe", + "size": 1902, + "annotations": { + "in-toto.io/predicate-type": "https://slsa.dev/provenance/v1" + } + } + ], + "subject": { + "mediaType": "application/vnd.oci.image.manifest.v1+json", + "digest": "sha256:fd2cd59f930cf2becfb547624bad96b2391338e765be2c76d5e16660a77ac598", + "size": 4684 + } +} \ No newline at end of file diff --git a/image/blobs/sha256/d447e3306f087f0a9b594fb8a92f46880560372c42349c9fc8c10e5bb634aadd b/image/blobs/sha256/d447e3306f087f0a9b594fb8a92f46880560372c42349c9fc8c10e5bb634aadd new file mode 100644 index 0000000000000000000000000000000000000000..6f0a2e08563125d93e0a43d5c0128584854e481c --- /dev/null +++ b/image/blobs/sha256/d447e3306f087f0a9b594fb8a92f46880560372c42349c9fc8c10e5bb634aadd @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:d447e3306f087f0a9b594fb8a92f46880560372c42349c9fc8c10e5bb634aadd +size 288195193 diff --git a/image/blobs/sha256/df5d19f5344d5c6d2ef185342673736620bb865afc93761139fdc7fdbac5614a b/image/blobs/sha256/df5d19f5344d5c6d2ef185342673736620bb865afc93761139fdc7fdbac5614a new file mode 100644 index 0000000000000000000000000000000000000000..d5a740ab0c6f0e813a62959e6bfc0cd9c4f7c4c0 Binary files /dev/null and b/image/blobs/sha256/df5d19f5344d5c6d2ef185342673736620bb865afc93761139fdc7fdbac5614a differ diff --git a/image/blobs/sha256/e1ff664083a6df132a0751cb98c17c425fd5908c3434270f38e2de56763f06ba b/image/blobs/sha256/e1ff664083a6df132a0751cb98c17c425fd5908c3434270f38e2de56763f06ba new file mode 100644 index 0000000000000000000000000000000000000000..b2e6aefd9ce6279b449a17323a5eb50ab128350d --- /dev/null +++ b/image/blobs/sha256/e1ff664083a6df132a0751cb98c17c425fd5908c3434270f38e2de56763f06ba @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:e1ff664083a6df132a0751cb98c17c425fd5908c3434270f38e2de56763f06ba +size 69483613 diff --git a/image/blobs/sha256/e324ba29d9a1eb6f52840139e2d929d488845bcadb738439b521f7900e1e22a3 b/image/blobs/sha256/e324ba29d9a1eb6f52840139e2d929d488845bcadb738439b521f7900e1e22a3 new file mode 100644 index 0000000000000000000000000000000000000000..a6af0bd11105811af2da8c9cb2a20a2870ecdd04 --- /dev/null +++ b/image/blobs/sha256/e324ba29d9a1eb6f52840139e2d929d488845bcadb738439b521f7900e1e22a3 @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:e324ba29d9a1eb6f52840139e2d929d488845bcadb738439b521f7900e1e22a3 +size 128074 diff --git a/image/blobs/sha256/e366a01215d904e93d01625e08bcb548e197f56606b17c5100bdc2ea9d89f907 b/image/blobs/sha256/e366a01215d904e93d01625e08bcb548e197f56606b17c5100bdc2ea9d89f907 new file mode 100644 index 0000000000000000000000000000000000000000..ad6aa44274ce3c1066e44e0b8e2b7abf1e4461d3 Binary files /dev/null and b/image/blobs/sha256/e366a01215d904e93d01625e08bcb548e197f56606b17c5100bdc2ea9d89f907 differ diff --git a/image/blobs/sha256/ed9f975b7148483c034ca6dec37793ef0d6ed4d60b99159a3dcd44fd56c847fe b/image/blobs/sha256/ed9f975b7148483c034ca6dec37793ef0d6ed4d60b99159a3dcd44fd56c847fe new file mode 100644 index 0000000000000000000000000000000000000000..6086c60ab89d6e5b4a9a105271d99681ac7ac485 --- /dev/null +++ b/image/blobs/sha256/ed9f975b7148483c034ca6dec37793ef0d6ed4d60b99159a3dcd44fd56c847fe @@ -0,0 +1 @@ +{"_type":"https://in-toto.io/Statement/v1","predicateType":"https://slsa.dev/provenance/v1","subject":[{"name":"pkg:docker/tt-model/diffusion-planner-p150@build-44d665005?platform=linux%2Famd64","digest":{"sha256":"fd2cd59f930cf2becfb547624bad96b2391338e765be2c76d5e16660a77ac598"}}],"predicate":{"buildDefinition":{"buildType":"https://github.com/moby/buildkit/blob/master/docs/attestations/slsa-definitions.md","resolvedDependencies":[{"uri":"pkg:docker/docker/dockerfile@1.7-labs","digest":{"sha256":"b99fecfe00268a8b556fad7d9c37ee25d716ae08a5d7320e6d51c4dd83246894"}},{"uri":"pkg:docker/ubuntu@22.04?platform=linux%2Famd64","digest":{"sha256":"5ec03bb3441e8b0bf3b4f9cd4629a1ae763010dc3035bb8da3ae6cf026486401"}},{"uri":"pkg:docker/ghcr.io/tenstorrent/tt-metal/tt-metalium/ubuntu-22.04-dev-amd64@latest?platform=linux%2Famd64","digest":{"sha256":"df9d279c7f85c17c6fad982d196802682d669cca1b7ced9cbaad8181339cd5fc"}}],"externalParameters":{"configSource":{"path":"Dockerfile"},"request":{"frontend":"gateway.v0","args":{"cmdline":"docker/dockerfile:1.7-labs","context:metalsrc":"local:metalsrc","frontend.caps":"moby.buildkit.frontend.contexts+forward","sharedkey:localdir:metalsrc":"metalsrc:4e7665ce329afc22","source":"docker/dockerfile:1.7-labs"},"locals":[{"name":"context"},{"name":"dockerfile"},{"name":"metalsrc"}],"root":{"configSource":{"path":"Dockerfile"},"request":{"args":{"context:metalsrc":"local:metalsrc","frontend.caps":"moby.buildkit.frontend.contexts+forward","sharedkey:localdir:metalsrc":"metalsrc:4e7665ce329afc22"}}},"compatibilityVersion":30}},"internalParameters":{"builderPlatform":"linux/amd64"}},"runDetails":{"builder":{"id":""},"metadata":{"invocationId":"39lol4fx5hun1bm90dy21s42z","startedOn":"2026-10-09T04:37:19.74554185Z","finishedOn":"2026-10-09T04:46:58.073923593Z","buildkit_metadata":{},"buildkit_completeness":{"request":false,"resolvedDependencies":false}}}}} \ No newline at end of file diff --git a/image/blobs/sha256/fc48d390b577ea739097ee841cf6e598acbbe40f0f42571c6cb450d57d7a09f2 b/image/blobs/sha256/fc48d390b577ea739097ee841cf6e598acbbe40f0f42571c6cb450d57d7a09f2 new file mode 100644 index 0000000000000000000000000000000000000000..55601b1066af3074c97efbc1e311129ce9a5c636 --- /dev/null +++ b/image/blobs/sha256/fc48d390b577ea739097ee841cf6e598acbbe40f0f42571c6cb450d57d7a09f2 @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:fc48d390b577ea739097ee841cf6e598acbbe40f0f42571c6cb450d57d7a09f2 +size 13065055 diff --git a/image/blobs/sha256/fd2cd59f930cf2becfb547624bad96b2391338e765be2c76d5e16660a77ac598 b/image/blobs/sha256/fd2cd59f930cf2becfb547624bad96b2391338e765be2c76d5e16660a77ac598 new file mode 100644 index 0000000000000000000000000000000000000000..638e2bb018255c5e6144d51310f441b80eb804d9 --- /dev/null +++ b/image/blobs/sha256/fd2cd59f930cf2becfb547624bad96b2391338e765be2c76d5e16660a77ac598 @@ -0,0 +1,126 @@ +{ + "schemaVersion": 2, + "mediaType": "application/vnd.oci.image.manifest.v1+json", + "config": { + "mediaType": "application/vnd.oci.image.config.v1+json", + "digest": "sha256:8bc4b6829641a7f2dde4ebe50b8189de860e7ec898b8570f0922a540b3c0ca6a", + "size": 15104 + }, + "layers": [ + { + "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", + "digest": "sha256:98c4455a98982b35380ec62f90d5cf0edff02eeb4b3bb18595d871fdeb0c0c67", + "size": 29751627 + }, + { + "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", + "digest": "sha256:e1ff664083a6df132a0751cb98c17c425fd5908c3434270f38e2de56763f06ba", + "size": 69483613 + }, + { + "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", + "digest": "sha256:df5d19f5344d5c6d2ef185342673736620bb865afc93761139fdc7fdbac5614a", + "size": 4396 + }, + { + "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", + "digest": "sha256:a15921f6aeba449853d8dfc7215c466ad00cff30329183f8b839522d33bc6552", + "size": 9565630 + }, + { + "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", + "digest": "sha256:be04a2af1cfe08908467229c0dd35fc629cb65ba98a497e7b627499f0c76821b", + "size": 138968702 + }, + { + "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", + "digest": "sha256:3ac9de4540f40bdf99f054641b2fb36a38f49253d1727fb08a73272d4208ce1b", + "size": 66024983 + }, + { + "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", + "digest": "sha256:d447e3306f087f0a9b594fb8a92f46880560372c42349c9fc8c10e5bb634aadd", + "size": 288195193 + }, + { + "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", + "digest": "sha256:18c4be700a3778a5b45eec8a3647e95f4376c403394168df7f0251356baad738", + "size": 115 + }, + { + "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", + "digest": "sha256:46f21edbef18d35cb9e2f400eeb6d8f8fac5873f8620e8d96829a9c01447ed52", + "size": 139083872 + }, + { + "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", + "digest": "sha256:ca0cef561ac457e6f5fe6e064c167bfc523b6e975437121969a8f39630a0a424", + "size": 41057991 + }, + { + "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", + "digest": "sha256:0936b6c31255cfe6d029dab00fc251d8771f5e3d228ac7da23e395e1161c6992", + "size": 41057719 + }, + { + "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", + "digest": "sha256:069395655fd57335d9426528f7f49731be6cf7ed963ca649e4ff7c9697d37e0a", + "size": 63869034 + }, + { + "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", + "digest": "sha256:fc48d390b577ea739097ee841cf6e598acbbe40f0f42571c6cb450d57d7a09f2", + "size": 13065055 + }, + { + "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", + "digest": "sha256:c74a07e1fb0aaf18cdccf3f41385a21dd137109c155363830e2ac2c565d7ac43", + "size": 1949486 + }, + { + "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", + "digest": "sha256:e366a01215d904e93d01625e08bcb548e197f56606b17c5100bdc2ea9d89f907", + "size": 6884 + }, + { + "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", + "digest": "sha256:c006f8f8bd1ac6c419a231b5d94521d87d03b141b76873d7a4d7fb68a933e57a", + "size": 848352 + }, + { + "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", + "digest": "sha256:4f4fb700ef54461cfa02571ae0db9a0dc1e0cdb5577484a6d75e68dc38e8acc1", + "size": 32 + }, + { + "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", + "digest": "sha256:4ad0ad64dc956c94bdcf581b56664b50f1d1ac6d091e6a7b1f28bf612bd7ab69", + "size": 1078 + }, + { + "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", + "digest": "sha256:a776bb9879e89222f1941fe93976727dbbfd5e8f5da8f65f42b05095ecb25808", + "size": 1169 + }, + { + "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", + "digest": "sha256:4f4fb700ef54461cfa02571ae0db9a0dc1e0cdb5577484a6d75e68dc38e8acc1", + "size": 32 + }, + { + "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", + "digest": "sha256:a897c24f6585a41459698990eabb275e716204a7310044da5cfda2b0f61809ee", + "size": 924 + }, + { + "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", + "digest": "sha256:e324ba29d9a1eb6f52840139e2d929d488845bcadb738439b521f7900e1e22a3", + "size": 128074 + }, + { + "mediaType": "application/vnd.oci.image.layer.v1.tar+gzip", + "digest": "sha256:ab2ce486a4537408cb227812400260a4fe3e3c19e860f0a2c65aae2a8d889f86", + "size": 4037 + } + ] +} \ No newline at end of file diff --git a/image/index.json b/image/index.json new file mode 100644 index 0000000000000000000000000000000000000000..7311f90043f3c44f5a0ff15a2309caa456abab2f --- /dev/null +++ b/image/index.json @@ -0,0 +1 @@ +{"schemaVersion":2,"mediaType":"application/vnd.oci.image.index.v1+json","manifests":[{"mediaType":"application/vnd.oci.image.index.v1+json","digest":"sha256:3b96d8ea71902fe6a00f1792dd41290839b7758070431464cd613cde3f6bf909","size":856,"annotations":{"io.containerd.image.name":"docker.io/tt-model/diffusion-planner-p150:3b96d8ea7190","org.opencontainers.image.ref.name":"3b96d8ea7190"}}]} \ No newline at end of file diff --git a/image/manifest.json b/image/manifest.json new file mode 100644 index 0000000000000000000000000000000000000000..de984727778571073cb0b159b2702d9a0849f5d4 --- /dev/null +++ b/image/manifest.json @@ -0,0 +1 @@ +[{"Config":"blobs/sha256/8bc4b6829641a7f2dde4ebe50b8189de860e7ec898b8570f0922a540b3c0ca6a","RepoTags":["tt-model/diffusion-planner-p150:3b96d8ea7190"],"Layers":["blobs/sha256/98c4455a98982b35380ec62f90d5cf0edff02eeb4b3bb18595d871fdeb0c0c67","blobs/sha256/e1ff664083a6df132a0751cb98c17c425fd5908c3434270f38e2de56763f06ba","blobs/sha256/df5d19f5344d5c6d2ef185342673736620bb865afc93761139fdc7fdbac5614a","blobs/sha256/a15921f6aeba449853d8dfc7215c466ad00cff30329183f8b839522d33bc6552","blobs/sha256/be04a2af1cfe08908467229c0dd35fc629cb65ba98a497e7b627499f0c76821b","blobs/sha256/3ac9de4540f40bdf99f054641b2fb36a38f49253d1727fb08a73272d4208ce1b","blobs/sha256/d447e3306f087f0a9b594fb8a92f46880560372c42349c9fc8c10e5bb634aadd","blobs/sha256/18c4be700a3778a5b45eec8a3647e95f4376c403394168df7f0251356baad738","blobs/sha256/46f21edbef18d35cb9e2f400eeb6d8f8fac5873f8620e8d96829a9c01447ed52","blobs/sha256/ca0cef561ac457e6f5fe6e064c167bfc523b6e975437121969a8f39630a0a424","blobs/sha256/0936b6c31255cfe6d029dab00fc251d8771f5e3d228ac7da23e395e1161c6992","blobs/sha256/069395655fd57335d9426528f7f49731be6cf7ed963ca649e4ff7c9697d37e0a","blobs/sha256/fc48d390b577ea739097ee841cf6e598acbbe40f0f42571c6cb450d57d7a09f2","blobs/sha256/c74a07e1fb0aaf18cdccf3f41385a21dd137109c155363830e2ac2c565d7ac43","blobs/sha256/e366a01215d904e93d01625e08bcb548e197f56606b17c5100bdc2ea9d89f907","blobs/sha256/c006f8f8bd1ac6c419a231b5d94521d87d03b141b76873d7a4d7fb68a933e57a","blobs/sha256/4f4fb700ef54461cfa02571ae0db9a0dc1e0cdb5577484a6d75e68dc38e8acc1","blobs/sha256/4ad0ad64dc956c94bdcf581b56664b50f1d1ac6d091e6a7b1f28bf612bd7ab69","blobs/sha256/a776bb9879e89222f1941fe93976727dbbfd5e8f5da8f65f42b05095ecb25808","blobs/sha256/4f4fb700ef54461cfa02571ae0db9a0dc1e0cdb5577484a6d75e68dc38e8acc1","blobs/sha256/a897c24f6585a41459698990eabb275e716204a7310044da5cfda2b0f61809ee","blobs/sha256/e324ba29d9a1eb6f52840139e2d929d488845bcadb738439b521f7900e1e22a3","blobs/sha256/ab2ce486a4537408cb227812400260a4fe3e3c19e860f0a2c65aae2a8d889f86"]}] \ No newline at end of file diff --git a/image/oci-layout b/image/oci-layout new file mode 100644 index 0000000000000000000000000000000000000000..1343d370fa7b18a594705346b415d647f611a1d1 --- /dev/null +++ b/image/oci-layout @@ -0,0 +1 @@ +{"imageLayoutVersion":"1.0.0"} \ No newline at end of file diff --git a/media/ATTRIBUTION.md b/media/ATTRIBUTION.md new file mode 100644 index 0000000000000000000000000000000000000000..3783308e24f4ade06cdc5217a6ede17a9e499795 --- /dev/null +++ b/media/ATTRIBUTION.md @@ -0,0 +1,55 @@ +# Media attribution: diffusion-planner-p150 demo renders + +Every render shows the **TT output**: the port running on one Tenstorrent Blackhole p150 (tt-nn; ETH dispatch, 1 +command queue, 12x10 grid; the pinned numerics of `tt-model.yaml`), called through the Python API +(`DiffusionPlanner.from_pretrained()`, then `model(inputs=...)`), which returns exactly what `POST /predict` returns. +The `*_tt_vs_cpu.png` views and the denoising view put it next to the fp32 CPU reference of the same Autoware network +(the port's torch reference, itself equal to ONNX Runtime on the shipped ONNX files, with the node's pre- and +post-processing), the oracle of the accuracy gates. + +What the bird's-eye views draw: the model's own input tensors in the current ego frame (lanes with their bounds, route +lanes, intersection polygons, stop lines, road borders, the 0.1 s neighbour histories), the 8 s ego plan (blue, a dot +every second), the predicted 8 s paths of the neighbours (thin lines in the class colour) and, for nuScenes, the +logged ego path (grey, hollow dots every second). The camera views draw the plan as a vehicle-width ribbon on the +ground plane of the CAM_FRONT key-frame image, and the logged path as a line. Renderer: the research renderer of the +public-data set (`research/diffusion-planner/public_data/scripts/dp_public_render.py`: same panels, palette, legend and +NC label as the CPU-reference renders of `research/diffusion-planner/media/`), driven by the workspace scripts +`logs/diffusion-planner/docs/scripts/{tt_demo_runs,render_tt_media}.py`; not shipped, because the dataset they read is +not. + +| file | source | license | +|---|---|---| +| `dp_kashiwanoha_dense_tt_vs_cpu.png` | the shipped sample `code/tt_diffusion_planner/samples/kashiwanoha_dense.npz`: a scripted scene on the Lanelet2 map AutowareFoundation/map-carla-kashiwanoha@0.2.0 | Apache-2.0 (map: Apache-2.0; scene and render: this repo) | +| `dp_kashiwanoha_dense_denoising_steps_tt.png` | the same sample: the ego row of the 11 solver iterates (`~/debug/denoising_steps`) | Apache-2.0 | +| `dp_straight_road_tt_vs_cpu.png` | the shipped sample `straight_road.npz`, a procedural scene of this repo | Apache-2.0 | +| `dp_tt_vs_cpu_agreement_99_scenes.png` | agreement statistics (ego displacement p150 vs CPU, per scene) of the 99 gated scenes: the 2 samples, 5 research scenes (kashiwanoha / CARLA Town10HD maps and procedural roads) and 92 planning instants converted from nuScenes v1.0-mini | a chart of this port's agreement numbers (no dataset content); nuScenes is credited in the image | +| `dp_nuscenes_scene-0061_kf06_bev_tt_NC.png` | nuScenes v1.0-mini scene-0061 (Singapore One-North, mini_train), key-frame 6 (t = 3.0 s): slip road before a left turn, following a van | **CC BY-NC-SA 4.0** + nuScenes Terms of Use (non-commercial) | +| `dp_nuscenes_scene-0061_kf18_bev_tt_NC.png`, `dp_nuscenes_scene-0061_kf18_cam_front_tt_NC.jpg` | scene-0061, key-frame 18 (t = 9.1 s), in the left turn; CAM_FRONT key-frame image | **CC BY-NC-SA 4.0** + nuScenes Terms of Use | +| `dp_nuscenes_scene-0103_kf12_bev_tt_NC.png`, `dp_nuscenes_scene-0103_kf12_cam_front_tt_NC.jpg` | nuScenes mini_val scene-0103 (Boston Seaport), key-frame 12 (t = 6.0 s) | **CC BY-NC-SA 4.0** + nuScenes Terms of Use | +| `dp_nuscenes_scene-0757_kf11_bev_tt_NC.png`, `dp_nuscenes_scene-0757_kf11_cam_front_tt_NC.jpg` | scene-0757 (Boston Seaport, mini_train), key-frame 11 (t = 5.3 s): approach to a signalised intersection with a bus crossing | **CC BY-NC-SA 4.0** + nuScenes Terms of Use | +| `dp_nuscenes_scene-0916_kf11_bev_tt_NC.png` | nuScenes mini_val scene-0916 (Singapore Queenstown, parking lot), key-frame 11 (t = 5.4 s) | **CC BY-NC-SA 4.0** + nuScenes Terms of Use | +| `dp_nuscenes_scene-0061_bev_tt_NC.gif` | scene-0061, every key-frame from t = 3.0 s to 19.2 s (2 Hz, 33 plans), shown at 2 frames/s | **CC BY-NC-SA 4.0** + nuScenes Terms of Use | + +What was changed in the nuScenes renders: the map expansion v1.3, CAN bus and annotations converted to the Autoware +Diffusion Planner input tensors (`research/diffusion-planner/public_data/scripts/nuscenes_dp.py`), bird's-eye views +rendered from those tensors and the p150 outputs; camera images resized from 1600x900 to 960x540 with the planned path +drawn and a caption band added. Every image is at most 960 px wide, nothing is cropped or zoomed in on people or +license plates, and the non-commercial label is drawn into every nuScenes image. + +## nuScenes (non-commercial) + +> Rendered from the nuScenes dataset (v1.0-mini, CAN bus expansion and map expansion v1.3), (c) Motional AD Inc., +> CC BY-NC-SA 4.0 and the nuScenes Terms of Use (https://www.nuscenes.org/terms-of-use). Non-commercial use only; +> adaptations under the same license. Motional does not endorse this work. Changes: converted to the Autoware Diffusion +> Planner input tensors, bird's-eye-view renders and resized camera images with the planned path drawn. Cite: +> H. Caesar et al., *nuScenes: A Multimodal Dataset for Autonomous Driving*, CVPR 2020. + +The `*_NC` files inherit CC BY-NC-SA 4.0: they are labelled "non-commercial" wherever they appear (each image's own +footer says so). No nuScenes data (images, tensors, outputs) is in this repository; only these renders are. + +## Kashiwanoha map (Apache-2.0) + +`code/tt_diffusion_planner/samples/kashiwanoha_dense.npz` and its renders are derived from the Lanelet2 map +AutowareFoundation/map-carla-kashiwanoha@0.2.0 (`lanelet2_map.osm`, Apache-2.0 per its dataset card; byte-identical to +the test map of autoware_universe `planning/autoware_diffusion_planner`), with an ego, a route and 88 agents scripted by +this port's research tools. `straight_road.npz` uses no external data. diff --git a/media/dp_kashiwanoha_dense_denoising_steps_tt.png b/media/dp_kashiwanoha_dense_denoising_steps_tt.png new file mode 100644 index 0000000000000000000000000000000000000000..f7739e0d586ec1aa459b0f18386ceefa02b2c899 Binary files /dev/null and b/media/dp_kashiwanoha_dense_denoising_steps_tt.png differ diff --git a/media/dp_kashiwanoha_dense_tt_vs_cpu.png b/media/dp_kashiwanoha_dense_tt_vs_cpu.png new file mode 100644 index 0000000000000000000000000000000000000000..1ff9e3472c010409faedbd6d45799e3649c631be --- /dev/null +++ b/media/dp_kashiwanoha_dense_tt_vs_cpu.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:30bdddea81eb6ca32ae9687d64f959911f38bde7ec254e98a3996bba2c260e48 +size 172970 diff --git a/media/dp_nuscenes_scene-0061_bev_tt_NC.gif b/media/dp_nuscenes_scene-0061_bev_tt_NC.gif new file mode 100644 index 0000000000000000000000000000000000000000..fef194b8c6c3814dd4d35ea90624f2da71c399fa --- /dev/null +++ b/media/dp_nuscenes_scene-0061_bev_tt_NC.gif @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:c0227fd8288d3379fc5f797e6b96a9bfc5167b9a220353bfe27b964f8e82d9d9 +size 1208869 diff --git a/media/dp_nuscenes_scene-0061_kf06_bev_tt_NC.png b/media/dp_nuscenes_scene-0061_kf06_bev_tt_NC.png new file mode 100644 index 0000000000000000000000000000000000000000..fa4665cec86764e2cfa9b8869e3a43e61306472d --- /dev/null +++ b/media/dp_nuscenes_scene-0061_kf06_bev_tt_NC.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:4b2ea8ca1064db08163b78f4efd4ceaf6424cb7d25d8c44a8f98e3d68b477f7c +size 195421 diff --git a/media/dp_nuscenes_scene-0061_kf18_bev_tt_NC.png b/media/dp_nuscenes_scene-0061_kf18_bev_tt_NC.png new file mode 100644 index 0000000000000000000000000000000000000000..fffd0a06694c53fbe926c81814053c6cae482d4b --- /dev/null +++ b/media/dp_nuscenes_scene-0061_kf18_bev_tt_NC.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:2d4bed37f90d7f912486b07d5adae51e7cc872fd9cacb5d62615c011f8a70032 +size 243157 diff --git a/media/dp_nuscenes_scene-0061_kf18_cam_front_tt_NC.jpg b/media/dp_nuscenes_scene-0061_kf18_cam_front_tt_NC.jpg new file mode 100644 index 0000000000000000000000000000000000000000..6700193cded36a5830815dc2d13fd0a5911ff8e8 --- /dev/null +++ b/media/dp_nuscenes_scene-0061_kf18_cam_front_tt_NC.jpg @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:13b9b21f5699682097f160c8dabf395a248c0edce14e3e2e31104843cde83932 +size 117295 diff --git a/media/dp_nuscenes_scene-0103_kf12_bev_tt_NC.png b/media/dp_nuscenes_scene-0103_kf12_bev_tt_NC.png new file mode 100644 index 0000000000000000000000000000000000000000..675c5758fa0dae4bebb2e6723d9407e94ac2ad52 --- /dev/null +++ b/media/dp_nuscenes_scene-0103_kf12_bev_tt_NC.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:ec9a0039e3a7a3e7577d944f8a1cfc29978f6a212ae6a5a80f9a5564aa30ff1a +size 156274 diff --git a/media/dp_nuscenes_scene-0103_kf12_cam_front_tt_NC.jpg b/media/dp_nuscenes_scene-0103_kf12_cam_front_tt_NC.jpg new file mode 100644 index 0000000000000000000000000000000000000000..f774f0ec6d3341b70a2ca34ecad14f85d100d5a0 --- /dev/null +++ b/media/dp_nuscenes_scene-0103_kf12_cam_front_tt_NC.jpg @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:48821161425d132265df27f8015aab537b2d9af6a8453af841abbb764b1e4812 +size 101401 diff --git a/media/dp_nuscenes_scene-0757_kf11_bev_tt_NC.png b/media/dp_nuscenes_scene-0757_kf11_bev_tt_NC.png new file mode 100644 index 0000000000000000000000000000000000000000..a66efa73b8b32ba25c0b119de53fae5b4347788a --- /dev/null +++ b/media/dp_nuscenes_scene-0757_kf11_bev_tt_NC.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:e61d27a3f622d918a90eb9895cd37619d64a7b2ff9cb88f332ba795940abc28e +size 165693 diff --git a/media/dp_nuscenes_scene-0757_kf11_cam_front_tt_NC.jpg b/media/dp_nuscenes_scene-0757_kf11_cam_front_tt_NC.jpg new file mode 100644 index 0000000000000000000000000000000000000000..c25687d40c9eecd69d01bc1f294d4078f29a383b Binary files /dev/null and b/media/dp_nuscenes_scene-0757_kf11_cam_front_tt_NC.jpg differ diff --git a/media/dp_nuscenes_scene-0916_kf11_bev_tt_NC.png b/media/dp_nuscenes_scene-0916_kf11_bev_tt_NC.png new file mode 100644 index 0000000000000000000000000000000000000000..28e042d297ac15e299ca0d085acdf90551dce396 --- /dev/null +++ b/media/dp_nuscenes_scene-0916_kf11_bev_tt_NC.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:af87b009cfe0e4da2f1893bfc18693efc969be2324425b3fc508f6396ae62ca7 +size 148756 diff --git a/media/dp_straight_road_tt_vs_cpu.png b/media/dp_straight_road_tt_vs_cpu.png new file mode 100644 index 0000000000000000000000000000000000000000..7b4a59bfb61283fb6b605618776805dd164b3644 --- /dev/null +++ b/media/dp_straight_road_tt_vs_cpu.png @@ -0,0 +1,3 @@ +version https://git-lfs.github.com/spec/v1 +oid sha256:8aed66d45342d403bf44e20c72be859c60d0602f586d281ae64f714d069a290e +size 115386 diff --git a/media/dp_tt_vs_cpu_agreement_99_scenes.png b/media/dp_tt_vs_cpu_agreement_99_scenes.png new file mode 100644 index 0000000000000000000000000000000000000000..5a19f73de06204f15cf2e7b71916a92301773446 Binary files /dev/null and b/media/dp_tt_vs_cpu_agreement_99_scenes.png differ diff --git a/patches/tt-metal-eth-dispatch.patch b/patches/tt-metal-eth-dispatch.patch new file mode 100644 index 0000000000000000000000000000000000000000..61fd39d1858b3e345c201c2a498423efd79dd753 --- /dev/null +++ b/patches/tt-metal-eth-dispatch.patch @@ -0,0 +1,84 @@ +diff --git a/tt_metal/hw/toolchain/main.ld b/tt_metal/hw/toolchain/main.ld +index 8cee67693d8..03732862192 100644 +--- a/tt_metal/hw/toolchain/main.ld ++++ b/tt_metal/hw/toolchain/main.ld +@@ -306,7 +306,8 @@ SECTIONS + LONG(TEXT_SIZE + #if defined(TYPE_KERNEL) && !(defined(COMPILE_FOR_NCRISC) && defined(ARCH_WORMHOLE)) \ + && !defined(COMPILE_FOR_AERISC) && !defined(COMPILE_FOR_SUBORDINATE_AERISC) \ +- && !defined(COMPILE_FOR_MAIN_AERISC) ++ && !defined(COMPILE_FOR_MAIN_AERISC) \ ++ && !(defined(ARCH_BLACKHOLE) && (defined(COMPILE_FOR_IERISC) || defined(COMPILE_FOR_SUBORDINATE_IERISC))) + - (__fw_export_text_end - TEXT_START) + #endif + ) +diff --git a/tt_metal/impl/dispatch/dispatch_query_manager.cpp b/tt_metal/impl/dispatch/dispatch_query_manager.cpp +index 38c31376178..fef7bafe6cd 100644 +--- a/tt_metal/impl/dispatch/dispatch_query_manager.cpp ++++ b/tt_metal/impl/dispatch/dispatch_query_manager.cpp +@@ -148,6 +148,15 @@ void DispatchQueryManager::reset(DispatchCoreConfig& dispatch_core_config, uint8 + (num_hw_cqs == 1 or dispatch_core_config_.get_dispatch_core_type() == DispatchCoreType::WORKER); + distributed_dispatcher_ = + (num_hw_cqs == 1 and dispatch_core_config_.get_dispatch_core_type() == DispatchCoreType::ETH); ++ // Blackhole ETH dispatch with 1 CQ needs 3 idle ETH cores (prefetch, dispatch, dispatch_s). ++ // Galaxy BH chips only have 2 idle ETH cores (the rest are trained links), so fall back to ++ // running without dispatch_s (dispatch_d sends go signals), which needs only 2 cores. ++ if (arch == tt::ARCH::BLACKHOLE and num_hw_cqs == 1 and ++ dispatch_core_config_.get_dispatch_core_type() == DispatchCoreType::ETH and ++ populate_all_logical_dispatch_cores(env_, num_hw_cqs_, dispatch_core_config_).size() < 3) { ++ dispatch_s_enabled_ = false; ++ distributed_dispatcher_ = false; ++ } + } + + go_signal_noc_ = (dispatch_s_enabled_ and arch != tt::ARCH::QUASAR) ? NOC::NOC_1 : NOC::NOC_0; +diff --git a/tt_metal/impl/dispatch/topology.cpp b/tt_metal/impl/dispatch/topology.cpp +index 3dc6377b652..813787a2bf5 100644 +--- a/tt_metal/impl/dispatch/topology.cpp ++++ b/tt_metal/impl/dispatch/topology.cpp +@@ -130,6 +130,12 @@ static const std::vector single_chip_arch_1cq = { + {2, 0, 0, 0, DISPATCH_S, {0, x, x, x}, {1, x, x, x}, k_dispatcher_s_noc}, + }; + ++// 1 CQ without dispatch_s (used for Blackhole ETH dispatch when only 2 idle ETH cores exist) ++static const std::vector single_chip_arch_1cq_no_dispatch_s = { ++ {0, 0, 0, 0, PREFETCH_HD, {x, x, x, x}, {1, x, x, x}, k_prefetcher_noc}, ++ {1, 0, 0, 0, DISPATCH_HD, {0, x, x, x}, {x, x, x, x}, k_dispatcher_noc}, ++}; ++ + static const std::vector single_chip_arch_2cq = { + {0, 0, 0, 0, PREFETCH_HD, {x, x, x, x}, {2, x, x, x}, k_prefetcher_noc}, + {1, 0, 0, 1, PREFETCH_HD, {x, x, x, x}, {3, x, x, x}, k_prefetcher_noc}, +@@ -464,7 +470,13 @@ std::vector DispatchTopology::generate_nodes( + auto populate_single_device = [&]() { + const bool is_quasar = this->descriptor_.cluster().arch() == tt::ARCH::QUASAR; + if (num_hw_cqs == 1) { +- return is_quasar ? quasar_single_chip_1cq : single_chip_arch_1cq; ++ if (is_quasar) { ++ return quasar_single_chip_1cq; ++ } ++ if (!this->get_dispatch_query_manager_().dispatch_s_enabled()) { ++ return single_chip_arch_1cq_no_dispatch_s; ++ } ++ return single_chip_arch_1cq; + } + if (is_quasar) { + return quasar_single_chip_2cq; +diff --git a/tt_metal/llrt/core_descriptor.cpp b/tt_metal/llrt/core_descriptor.cpp +index 0ada4eff2f0..1a2358d6cf7 100644 +--- a/tt_metal/llrt/core_descriptor.cpp ++++ b/tt_metal/llrt/core_descriptor.cpp +@@ -290,6 +290,13 @@ const tt::core_descriptor_t& MetalEnvImpl::get_core_descriptor_config( + coord = RelativeCoreCoord({.x = core_node[0].as(), .y = core_node[1].as()}); + if (get_core_type_from_config(dispatch_core_config) == CoreType::ETH) { + auto logical_coord = get_core_coord_from_relative(coord, grid_size); ++ // Skip ETH dispatch entries that do not exist on this chip (harvested ETH cores, e.g. ++ // Blackhole Galaxy chips expose fewer than the 14 ETH cores listed in the YAML). ++ const size_t num_logical_eth_cores = ++ get_cluster().get_soc_desc(device_id).get_cores(CoreType::ETH, CoordSystem::LOGICAL).size(); ++ if (logical_coord.y >= num_logical_eth_cores) { ++ continue; ++ } + if (logical_active_eth_cores.contains(logical_coord)) { + continue; + } diff --git a/pyproject.toml b/pyproject.toml new file mode 100644 index 0000000000000000000000000000000000000000..64de3d1537c7a078013e9661c5270ff79d021218 --- /dev/null +++ b/pyproject.toml @@ -0,0 +1,60 @@ +# SPDX-License-Identifier: Apache-2.0 +# Python package of diffusion-planner-p150, from the ROOT of the model repo: +# pip install -e . # the Python API +# pip install -e ".[server,test]" # + the HTTP server and the tests +# +# Install it into an environment that already has tt-metal's ttnn (and its torch 2.11.0+cpu): ttnn and torch are +# NOT declared, so pip never replaces the tt-metal build's own torch. The container image installs the same set +# through tt-model.yaml runtime.packages. +# +# This file lives at the repo root ON PURPOSE and must never move into code/: tt-model copies code/ over /opt/tt-metal +# right before `uv pip install -e /opt/tt-metal` in the image build, so a code/pyproject.toml would replace tt-metal's +# own (project ttnn) and break `import ttnn` in the image (research/PACKAGING_PILOT.md problem 1; +# research/packaging/scripts/check_bundle.py refuses it). It is not in tt-model.yaml extra_code: it reaches the Hub as +# an overlay file next to README.md (SERVING.md section 2), like the other top-level docs. +[build-system] +requires = ["setuptools>=77"] +build-backend = "setuptools.build_meta" + +[project] +name = "tt-diffusion-planner" +dynamic = ["version"] # tt_diffusion_planner.__version__, the version /openapi.json reports +description = "Diffusion Planner v5.0 (Autoware diffusion_planner) on Tenstorrent Blackhole p150 (tt-nn), with a Python API and an HTTP server" +readme = "code/PYTHON.md" +license = "Apache-2.0" +requires-python = ">=3.10" +dependencies = [ + "numpy>=1.24.4,<2", # ttnn needs numpy 1.x + "pillow", + "pyyaml", # Autoware *.param.yaml + "huggingface_hub", # weights pointer: AutowareFoundation/diffusion_planner@423efde67f5414734da43a7ad856c17ceb8b51aa + "safetensors", + "onnx>=1.17,<2", # reads the ONNX initializers (data only) +] + +[project.optional-dependencies] +server = ["fastapi", "uvicorn", "pydantic>=2"] +test = ["pytest", "httpx"] +reference = ["onnxruntime"] # ONNX Runtime golden (research / tests only, not in the image) + +[project.urls] +"Model card" = "https://huggingface.co/changh95/diffusion-planner-p150" +"Weights" = "https://huggingface.co/AutowareFoundation/diffusion_planner" +"Autoware package" = "https://github.com/autowarefoundation/autoware_universe/tree/9ceaccf026c31ffc5319bc9eeb4bd7bede0af3fd/planning/autoware_diffusion_planner" + +[tool.setuptools.dynamic] +version = { attr = "tt_diffusion_planner.__version__" } + +[tool.setuptools.packages.find] +where = ["code"] # import root (code/conftest.py and code/scripts/ are not packaged) +include = ["tt_diffusion_planner*"] # with the vendored tt_diffusion_planner.ttaw (common/tools/vendor.py; never edit it) +exclude = ["tt_diffusion_planner.tests*"] + +[tool.setuptools.package-data] +"tt_diffusion_planner" = ["samples/*", "calib/*.json"] +"tt_diffusion_planner.tt" = ["kernels/*.cpp", "kernels/*.hpp", "kernels/*.h"] +"tt_diffusion_planner.ttaw" = ["API.md", "VENDORED.json"] +"tt_diffusion_planner.ttaw.ops" = ["kernels/*.cpp", "kernels/*.hpp", "kernels/*.h"] + +[tool.pytest.ini_options] +markers = ["device: needs a Tenstorrent chip (skipped on host-only runs)"] diff --git a/requirements.lock b/requirements.lock new file mode 100644 index 0000000000000000000000000000000000000000..7bbc9545180e97f96733fea535131952f67d6890 --- /dev/null +++ b/requirements.lock @@ -0,0 +1,80 @@ +annotated-doc==0.0.5 +annotated-types==0.8.0 +anyio==4.15.1 +certifi==2026.7.22 +cfgv==3.5.0 +charset-normalizer==3.5.2 +click==8.5.0 +contourpy==1.3.3 +cycler==0.12.1 +distlib==0.4.3 +distro==1.9.0 +elastic-transport==9.4.2 +elasticsearch==9.5.1 +fastapi==0.143.0 +filelock==4.0.12 +fonttools==4.66.1 +fsspec==2026.9.0 +graphviz==0.21 +h11==0.16.0 +hf-xet==1.7.0 +httpcore2==2.13.1 +httpx2==2.13.1 +huggingface_hub==2.2.0 +identify==2.6.20 +idna==3.20 +Jinja2==3.1.6 +kiwisolver==1.5.1 +linkify-it-py==2.2.0 +loguru==0.7.3 +markdown-it-py==4.2.0 +MarkupSafe==3.0.4 +matplotlib==3.11.2 +mdit-py-plugins==0.6.1 +mdurl==0.1.2 +ml_dtypes==0.5.4 +mpmath==1.3.0 +networkx==3.7 +nodeenv==1.11.0 +numpy==1.26.4 +onnx==1.23.2 +opentelemetry-api==1.45.1 +packaging==26.3 +pandas==3.0.6 +pillow==12.3.0 +platformdirs==4.12.4 +pre_commit==4.6.2 +protobuf==7.36.2 +psutil==7.2.2 +pydantic==2.14.0 +pydantic_core==2.50.0 +Pygments==2.21.0 +pyluwen==0.10.0 +pyparsing==3.3.3 +python-dateutil==2.9.0.post0 +python-discovery==1.6.1 +PyYAML==6.0.3 +requests==2.34.2 +rich==15.0.0 +safetensors==0.8.0 +seaborn==0.13.2 +setuptools==80.10.2 +setuptools-scm==8.1.0 +six==1.17.0 +sniffio==1.3.1 +starlette==1.7.0 +sympy==1.14.0 +textual==8.2.8 +tomli==2.5.0 +torch==2.11.0+cpu +tqdm==4.70.1 +truststore==0.10.4 +tt-smi==6.7.0 +tt-tools-common==1.6.0 +tt-umd==0.9.12 +typing-inspection==0.4.4 +typing_extensions==4.16.0 +urllib3==2.8.0 +uvicorn==0.54.0 +virtualenv==21.14.6 +wheel==0.48.0 diff --git a/tt-model.yaml b/tt-model.yaml new file mode 100644 index 0000000000000000000000000000000000000000..3fbeb85108c6f6b2a7bca8c7a986784ad933dbdf --- /dev/null +++ b/tt-model.yaml @@ -0,0 +1,172 @@ +# SPDX-License-Identifier: Apache-2.0 +# tt-model-manager container manifest (schema 5.1): Diffusion Planner v5.0 (Autoware diffusion_planner) on one Blackhole p150. +# Conventions: research/BUNDLE_CONVENTIONS.md (workspace) -- section 3 explains every field. +# +# Run EVERY tt-model command FROM THIS DIRECTORY: source.tt_metal and extra_code[].root resolve against the +# process CWD, not this file; --out must point OUTSIDE this directory (package deletes / first). +# python -c "from tt_kernel.container_manifest import load_container_manifest as L; L('tt-model.yaml', check_sources=True)" +# sg docker -c "tt-model package --container tt-model.yaml --out /home/ubuntu/experiments/tt-models/build" +# (serve + smoke + stop under bin/devrun, once per serve profile: code/scripts/container_smoke.sh, SERVING.md 2) +# HF_TOKEN= tt-model push /home/ubuntu/experiments/tt-models/build/diffusion-planner-p150 --publish +schema: "5.1" + +repo: "changh95/diffusion-planner-p150" +name: "diffusion-planner-p150" + +# A POINTER to the network Autoware deploys (ansible artifacts role: diffusion_planner @ v5.0), never baked into +# the image. Pinned to the COMMIT the tag points to -- HfApi().model_info(repo, revision="v5.0").sha, not +# list_repo_refs().tags[].target_commit (that is the annotated-tag object) -- because tags can move and `package` +# does not resolve weights refs. `serve` / `pull --with-weights` pre-download exactly these files at this sha; the +# launcher exports the sha as TT_MODEL_WEIGHTS_REVISION (and serve.env repeats it as TT_WEIGHTS_REVISION). +weights: + repo: "AutowareFoundation/diffusion_planner" + revision: "423efde67f5414734da43a7ad856c17ceb8b51aa" + allow_patterns: ["diffusion_planner_encoder.onnx", "diffusion_planner_decoder.onnx", "diffusion_planner_turn_indicator.onnx", "diffusion_planner.param.json"] + +kind: tt-dit-server # the model's own FastAPI app under uvicorn (no vLLM); see the conventions doc +arch: blackhole + +source: + # tt-metal main 44d6650 (v0.80.0-dev20261006) WITH patches/tt-metal-eth-dispatch.patch applied: a dirty tree, + # packaged as-is (the image is built from it, so it carries ETH dispatch). torch 2.11.0 pin read from it. + tt_metal: "/home/ubuntu/experiments/tt-models/tt-metal" + # Schema minimum (>= 1 path relative to the tt-metal tree). The port imports nothing from tt-metal's models/. + code: + - models/common/lightweightmodule.py + # The port itself. code/ on the Hub becomes EXACTLY this set (+ the file above) at every push, and it is + # COPY'd over /opt/tt-metal/ in the image (PYTHONPATH=/opt/tt-metal) BEFORE tt-metal's C++ build and its + # `uv pip install -e /opt/tt-metal`. So never list a file that tt-metal's root also has and its build reads + # (pyproject.toml, setup.py, setup.cfg, MANIFEST.in, README.md, uv.toml, CMakeLists.txt, build_metal.sh, ...) or + # anything under its build trees (tt_metal/, ttnn/, cmake/, tools/, ...): check_bundle.py refuses them. The pip + # project is the repo-ROOT pyproject.toml (`pip install -e .`), published as an overlay file, never from here. + # conftest.py is deliberate: it replaces tt-metal's root conftest.py in the builder stage, which no build step + # reads, and the runtime stage never copies tt-metal's, so in the image it is the only root conftest. + extra_code: + - root: code + paths: + - "tt_diffusion_planner" + - scripts + - conftest.py + - PYTHON.md + ubuntu: "22.04" + python: "3.12" + +runtime: + app: "tt_diffusion_planner.server.app:app" + mesh_shape_env: TT_MESH_SHAPE + # On top of the kind defaults (fastapi, uvicorn, pydantic>=2, pillow) and the auto-pinned torch==2.11.0+cpu. + # numpy stays < 2 (ttnn). Add per model: scipy, opencv-python-headless>=4.9,<4.12, ... + packages: + - "numpy>=1.24.4,<2" + - huggingface_hub + - safetensors + - pyyaml + - "onnx>=1.17,<2" + # After the first green build: copy /diffusion-planner-p150/requirements.lock next to this file, check it, and + # uncomment (resolved relative to THIS file): + # lock: requirements.lock + +serve: + port: 20000 + hardware: p150 + mesh_device: P150 + env: + TT_WEIGHTS_REVISION: "423efde67f5414734da43a7ad856c17ceb8b51aa" # older tt-model exports no revision; the app also reads TT_MODEL_WEIGHTS_REVISION + TT_METAL_VISIBLE_DEVICES: "0" + # The weights repos are public and ungated: never send a token. serve mounts the host HF cache (which may hold a + # stored token file) at /hf, so tell huggingface_hub not to pick one up implicitly (PACKAGING_PILOT.md problem 10). + HF_HUB_DISABLE_IMPLICIT_TOKEN: "1" + # Pinned dispatch: ETH, the 12x10 grid every published number uses (D14). Profiles must not override it: + # check_bundle.py refuses any *_DISPATCH other than eth, and the container smoke fails unless /info reports + # dispatch=eth, grid=12x10 and these pins. + "DIFFUSION_PLANNER_DISPATCH": "eth" + "DIFFUSION_PLANNER_NUM_CQS": "1" + "DIFFUSION_PLANNER_VARIANT": "default" + # Pinned numerics: the configuration every gate and every number of the card was measured with (PORT_LOG decision + # 11; tt/config.py KNOBS, one value per knob = KNOBS.serve_env(), checked by tests/test_bundle_host.py). An empty + # value is the knob's empty default; an empty PRECISION keeps tt/config.py DEFAULT_PRECISION. Changing any of them + # changes the numerics, so the gates must be re-run first (OPT_BASELINE.md "What the numerics defaults cost"). + "DIFFUSION_PLANNER_LN_FP32": "enc.mixer.*,dec.*" + "DIFFUSION_PLANNER_HIDDEN_FP32": "" + "DIFFUSION_PLANNER_SPLIT_MATMUL": "enc.island.*,enc.pre.*,dec.*" + "DIFFUSION_PLANNER_ATTN_FP32_ACC": "" + "DIFFUSION_PLANNER_ATTN_MATMUL": "enc.fusion.attn,dec.*" + "DIFFUSION_PLANNER_PRECISION": "" + +# No serve_profiles: one profile (the default). v5.0 is the only graph the current node loads (the v3.x / v4.0 tags +# hold the same file names and cannot be profiles of one bundle; BUNDLE_CONVENTIONS.md section 13), and the numerics +# configuration is pinned above, so there is no load-time variant to offer. + +# Build-time assertions run INSIDE the finished image as uid 1000, no device, no weights. +verify: + # the image must carry the ETH-dispatch patch (12x10 grid): fail the build otherwise (PLAN.md 6.2 step 1) + - "assert 'single_chip_arch_1cq_no_dispatch_s' in open('/opt/tt-metal/tt_metal/impl/dispatch/topology.cpp').read()" + - "from tt_diffusion_planner.ttaw.device import eth_dispatch_patch_present as p; assert p(), 'ttaw does not see the ETH patch'" + - "import tt_diffusion_planner.server.app as a; assert a.app" + - "from tt_diffusion_planner.server.app import parse_mesh_shape as p; assert p('1x1') == p('(1, 1)') == p('1,1') == (1, 1)" + - "import numpy; assert int(numpy.__version__.split('.')[0]) < 2, numpy.__version__" + - "import onnx, yaml, huggingface_hub, safetensors, PIL" + - "from tt_diffusion_planner.api import DiffusionPlanner; assert DiffusionPlanner.DEFAULT_REVISION == '423efde67f5414734da43a7ad856c17ceb8b51aa'" + - "import ttnn; assert hasattr(ttnn, 'DispatchCoreConfig') and hasattr(ttnn, 'begin_trace_capture') and hasattr(ttnn, 'execute_trace') and hasattr(ttnn, 'release_trace')" + # the shipped sample and its stored CPU-reference output: the container smoke compares /predict with it + - "from pathlib import Path; d = Path('/opt/tt-metal/tt_diffusion_planner/samples'); assert all((d / f).is_file() for f in ('kashiwanoha_dense.npz', 'kashiwanoha_dense.reference.json'))" + # no custom kernel ships yet (code/tt_diffusion_planner/tt/kernels/README.md); add one line per kernel file when one does + +card: + # Lede under the title of the GENERATED card (tt-model package). The published README.md is the hand-written + # card of this directory (overlaid before push); keep the facts here and there identical. + description: > + Diffusion Planner v5.0 (Autoware diffusion_planner): the network Autoware deploys (autoware_diffusion_planner, weights AutowareFoundation/diffusion_planner + v5.0) running on one Tenstorrent Blackhole p150 via tt-nn, the whole plan (scene encoder, 11 DiT evaluations of the DPM-Solver++(2M) loop, turn-indicator head) as one metal trace: + the Autoware planner tensors (ego and neighbour histories, lanes, route, polygons, line strings, goal, ego shape, turn-indicator history) in, an 8 s ego trajectory, predicted paths of the valid neighbours and a turn-indicator command out. + architecture: "Six MLP-Mixer entity encoders + small MLP encoders, a 6-block transformer fusion encoder (564 scene tokens), a 3-block DiT decoder with adaLN (self-attention over 321 agents, cross-attention to the scene) evaluated 11 times by a DPM-Solver++(2M) loop, and a linear turn-indicator head (three ONNX graphs, 14.55 M parameters)" + status: > + First release, baseline port (optimization pending). Community port, validated by agreement with the fp32 CPU + reference on 99 scenes (2 shipped samples, 5 research scenes, 92 nuScenes v1.0-mini planning instants). + intended_use: > + Running and benchmarking the network of Autoware's diffusion planner node on one Blackhole p150, fed with the planner tensors the node builds from ROS messages and the Lanelet2 map (that conversion stays with the client). + out_of_scope_use: > + Safety-critical driving decisions (this is not a certified Autoware component); multi-chip meshes; batch > 1; + inputs outside the Autoware sensor configuration the weights were trained for. Closed-loop vehicle control; the node's guidance services (start / stop / centerline guidance). + quickstart: | + The first start compiles the kernels and captures the trace (minutes on a cold cache, seconds later); the + server is ready when its log shows `Application startup complete`. Build a request from a file with + `python3 code/tt_diffusion_planner/server/client.py --inputs --out req.json`, then + `curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json`. + usage: | + `POST /predict` (JSON): `inputs` (the 15 raw tensors of the Autoware node's `create_input_data()` in the ego frame, before normalization: a base64 `.npz` or `{"format": "json", "arrays": {...}}`); optional `params` (`velocity_smoothing_window` 8, `stopping_threshold` 0.3, `turn_indicator_keep_offset` -1.25, `return_denoising_steps` false), `output_format`. `GET /health`, `GET /info`, `GET /v1/models` (stub). + Full request / response contract: `SERVING.md` section 3. + performance: | + | metric | value | + |---|---| + | back-to-back trace replays (device time per plan) | 102.04 ms (9.80 plans/s) | + | one blocking plan (device trace: encoder + 11 DiT evaluations + solver + turn head) | 102.13 ms | + | Python `model()` call, shipped sample | 117.9 ms p50 | + | served `/predict` `timing_ms.total` (uvicorn on the host) | 123.1 ms median | + | agreement with the fp32 CPU reference, 99 scenes | ego max 0.313 m / mean 0.143 m worst case (gates 1.0 / 0.3 m), turn command 99 / 99 | + | configuration | ETH dispatch, 1 CQ, 12x10 grid, one p150; the pinned numerics | + limitations: | + - First release: baseline port, optimization pending. The numerics the end-to-end gates need cost 33.4 ms per plan (the first optimization target); the encoder's 4-8-core channel matmuls another 25.8 ms (OPT_REPORT.md). + - Stateless API: the node's turn-indicator hold window, the RTC prefix / temperature of the initial solver state and the agent / ego histories are the client's; converting ROS messages and the Lanelet2 map into the 15 input tensors is the client's; guidance services off. + - Batch 1, one plan per request; the v5.0 capacities (320 neighbours, 140 lanes, 25 route lanes, 10 polygons, 60 line strings) and 10 DPM-Solver steps are compiled into the trace. + - Accuracy is agreement with the fp32 CPU reference (99 scenes); no dataset-level planning metric: the paper's nuPlan closed-loop benchmark was not run, and the weights were trained on non-public TIER IV data. + - ETH dispatch, single chip: no multi-chip mesh; WORKER dispatch (11x10) is an A/B opt-in. + - Not a certified Autoware component; not for safety-critical driving decisions. + risks: > + Device numerics (split hi / lo matmuls, fp32 LayerNorm and attention, bf16 weights elsewhere) move the plan + slightly from the fp32 CPU reference: up to 0.31 m (max) / 0.14 m (mean over the 8 s) on the worst of 99 scenes, + a few cm on most; the sensitivity is chaotic per scene, so any numerics change must be re-checked on all scenes. + The driving quality is the weights' (trained by TIER IV on non-public data): on nuScenes, a domain they never + saw, the plans are plausible but conservative. A client that skips the node's state (turn-indicator hold window, + agent and ego histories, RTC prefix) gets different behaviour than the Autoware node. + licensing: | + - Weights: [AutowareFoundation/diffusion_planner](https://huggingface.co/AutowareFoundation/diffusion_planner) at tag `v5.0` (commit `423efde67f5414734da43a7ad856c17ceb8b51aa`), Apache-2.0 per its model card; not redistributed here. The upstream card states that TIER IV trained the models on TIER IV synthetic and real driving data; the dataset composition is not publicly documented. + - Pre-/post-processing ported from autoware_universe `planning/autoware_diffusion_planner` @ `9ceaccf026c31ffc5319bc9eeb4bd7bede0af3fd` (Apache-2.0). + - Port and serving code (`code/`): Apache-2.0. + related: | + Other Autoware models on Blackhole: the "Autoware" collection of https://huggingface.co/changh95 + license: + id: apache-2.0 + pipeline_tag: "robotics" + base_model: + - "AutowareFoundation/diffusion_planner" diff --git a/tt_kernel_manifest.json b/tt_kernel_manifest.json new file mode 100644 index 0000000000000000000000000000000000000000..aa50edab47148b685443d1372244fc10ee47efb3 --- /dev/null +++ b/tt_kernel_manifest.json @@ -0,0 +1,149 @@ +{ + "schema_version": "5.1", + "name": "diffusion-planner-p150", + "tt_metal_version": "0.65.2.dev11169+g44d66500520", + "arch": "blackhole", + "device_count": 1, + "producer": { + "tt_kernel_version": "0.1.0", + "created_at": "2026-10-09T04:47:53.805841+00:00", + "hostname": "hchang-bh" + }, + "weights": { + "repo_id": "AutowareFoundation/diffusion_planner", + "revision": "423efde67f5414734da43a7ad856c17ceb8b51aa", + "allow_patterns": [ + "diffusion_planner_encoder.onnx", + "diffusion_planner_decoder.onnx", + "diffusion_planner_turn_indicator.onnx", + "diffusion_planner.param.json" + ], + "ignore_patterns": null, + "repo_type": "model" + }, + "mesh": null, + "entrypoint": null, + "resources": null, + "capabilities": null, + "env": {}, + "bundled": null, + "deps": null, + "container": { + "image": { + "registry": "hf", + "repository": "diffusion-planner-p150", + "tag": "tt-model/diffusion-planner-p150:3b96d8ea7190", + "digest": "sha256:3b96d8ea71902fe6a00f1792dd41290839b7758070431464cd613cde3f6bf909" + }, + "kind": "tt-dit-server", + "runtime": { + "app": "tt_diffusion_planner.server.app:app", + "mesh_shape_env": "TT_MESH_SHAPE", + "packages": [ + "numpy>=1.24.4,<2", + "huggingface_hub", + "safetensors", + "pyyaml", + "onnx>=1.17,<2" + ], + "lock": "requirements.lock" + }, + "serve": { + "hardware": "p150", + "mesh_device": "P150", + "port": 20000, + "max_model_len": null, + "max_num_seqs": null, + "block_size": null, + "server_timeout": null, + "capabilities": null, + "additional_config": {}, + "args": [], + "env": { + "TT_WEIGHTS_REVISION": "423efde67f5414734da43a7ad856c17ceb8b51aa", + "TT_METAL_VISIBLE_DEVICES": "0", + "HF_HUB_DISABLE_IMPLICIT_TOKEN": "1", + "DIFFUSION_PLANNER_DISPATCH": "eth", + "DIFFUSION_PLANNER_NUM_CQS": "1", + "DIFFUSION_PLANNER_VARIANT": "default", + "DIFFUSION_PLANNER_LN_FP32": "enc.mixer.*,dec.*", + "DIFFUSION_PLANNER_HIDDEN_FP32": "", + "DIFFUSION_PLANNER_SPLIT_MATMUL": "enc.island.*,enc.pre.*,dec.*", + "DIFFUSION_PLANNER_ATTN_FP32_ACC": "", + "DIFFUSION_PLANNER_ATTN_MATMUL": "enc.fusion.attn,dec.*", + "DIFFUSION_PLANNER_PRECISION": "" + } + }, + "serve_profiles": [ + { + "hardware": null, + "mesh_device": null, + "port": null, + "max_model_len": null, + "max_num_seqs": null, + "block_size": null, + "server_timeout": null, + "capabilities": null, + "additional_config": {}, + "args": [], + "env": {}, + "name": "default", + "description": null + } + ], + "default_profile": null, + "code_dir": "code", + "verify": [ + "assert 'single_chip_arch_1cq_no_dispatch_s' in open('/opt/tt-metal/tt_metal/impl/dispatch/topology.cpp').read()", + "from tt_diffusion_planner.ttaw.device import eth_dispatch_patch_present as p; assert p(), 'ttaw does not see the ETH patch'", + "import tt_diffusion_planner.server.app as a; assert a.app", + "from tt_diffusion_planner.server.app import parse_mesh_shape as p; assert p('1x1') == p('(1, 1)') == p('1,1') == (1, 1)", + "import numpy; assert int(numpy.__version__.split('.')[0]) < 2, numpy.__version__", + "import onnx, yaml, huggingface_hub, safetensors, PIL", + "from tt_diffusion_planner.api import DiffusionPlanner; assert DiffusionPlanner.DEFAULT_REVISION == '423efde67f5414734da43a7ad856c17ceb8b51aa'", + "import ttnn; assert hasattr(ttnn, 'DispatchCoreConfig') and hasattr(ttnn, 'begin_trace_capture') and hasattr(ttnn, 'execute_trace') and hasattr(ttnn, 'release_trace')", + "from pathlib import Path; d = Path('/opt/tt-metal/tt_diffusion_planner/samples'); assert all((d / f).is_file() for f in ('kashiwanoha_dense.npz', 'kashiwanoha_dense.reference.json'))" + ], + "built": { + "image": "tt-model/diffusion-planner-p150:3b96d8ea7190", + "repo": "changh95/diffusion-planner-p150", + "tt_model_version": "0.1.0", + "created_at": "2026-10-09T04:37:19+00:00", + "tt_metal": { + "sha": "44d66500520fda9f2c7060c0f6b41ec48f7ab37e", + "describe": "v0.80.0-dev20261006-78-g44d6650052-dirty", + "dirty": true, + "scm_version": "0.65.2.dev11169+g44d66500520", + "mode": "local", + "remote": "https://github.com/tenstorrent/tt-metal.git", + "branch": "main", + "pushed": true + }, + "code_sha256": "c0e7abb7888098a9319ab5c66a10c4fd4009fc1542326b26507554c04f05a481", + "image_digest": "sha256:3b96d8ea71902fe6a00f1792dd41290839b7758070431464cd613cde3f6bf909" + }, + "card": { + "description": "Diffusion Planner v5.0 (Autoware diffusion_planner): the network Autoware deploys (autoware_diffusion_planner, weights AutowareFoundation/diffusion_planner v5.0) running on one Tenstorrent Blackhole p150 via tt-nn, the whole plan (scene encoder, 11 DiT evaluations of the DPM-Solver++(2M) loop, turn-indicator head) as one metal trace: the Autoware planner tensors (ego and neighbour histories, lanes, route, polygons, line strings, goal, ego shape, turn-indicator history) in, an 8 s ego trajectory, predicted paths of the valid neighbours and a turn-indicator command out.\n", + "quickstart": "The first start compiles the kernels and captures the trace (minutes on a cold cache, seconds later); the\nserver is ready when its log shows `Application startup complete`. Build a request from a file with\n`python3 code/tt_diffusion_planner/server/client.py --inputs --out req.json`, then\n`curl -s localhost:20000/predict -H 'Content-Type: application/json' -d @req.json`.\n", + "architecture": "Six MLP-Mixer entity encoders + small MLP encoders, a 6-block transformer fusion encoder (564 scene tokens), a 3-block DiT decoder with adaLN (self-attention over 321 agents, cross-attention to the scene) evaluated 11 times by a DPM-Solver++(2M) loop, and a linear turn-indicator head (three ONNX graphs, 14.55 M parameters)", + "status": "First release, baseline port (optimization pending). Community port, validated by agreement with the fp32 CPU reference on 99 scenes (2 shipped samples, 5 research scenes, 92 nuScenes v1.0-mini planning instants).\n", + "intended_use": "Running and benchmarking the network of Autoware's diffusion planner node on one Blackhole p150, fed with the planner tensors the node builds from ROS messages and the Lanelet2 map (that conversion stays with the client).\n", + "out_of_scope_use": "Safety-critical driving decisions (this is not a certified Autoware component); multi-chip meshes; batch > 1; inputs outside the Autoware sensor configuration the weights were trained for. Closed-loop vehicle control; the node's guidance services (start / stop / centerline guidance).\n", + "usage": "`POST /predict` (JSON): `inputs` (the 15 raw tensors of the Autoware node's `create_input_data()` in the ego frame, before normalization: a base64 `.npz` or `{\"format\": \"json\", \"arrays\": {...}}`); optional `params` (`velocity_smoothing_window` 8, `stopping_threshold` 0.3, `turn_indicator_keep_offset` -1.25, `return_denoising_steps` false), `output_format`. `GET /health`, `GET /info`, `GET /v1/models` (stub).\nFull request / response contract: `SERVING.md` section 3.\n", + "performance": "| metric | value |\n|---|---|\n| back-to-back trace replays (device time per plan) | 102.04 ms (9.80 plans/s) |\n| one blocking plan (device trace: encoder + 11 DiT evaluations + solver + turn head) | 102.13 ms |\n| Python `model()` call, shipped sample | 117.9 ms p50 |\n| served `/predict` `timing_ms.total` (uvicorn on the host) | 123.1 ms median |\n| agreement with the fp32 CPU reference, 99 scenes | ego max 0.313 m / mean 0.143 m worst case (gates 1.0 / 0.3 m), turn command 99 / 99 |\n| configuration | ETH dispatch, 1 CQ, 12x10 grid, one p150; the pinned numerics |\n", + "limitations": "- First release: baseline port, optimization pending. The numerics the end-to-end gates need cost 33.4 ms per plan (the first optimization target); the encoder's 4-8-core channel matmuls another 25.8 ms (OPT_REPORT.md).\n- Stateless API: the node's turn-indicator hold window, the RTC prefix / temperature of the initial solver state and the agent / ego histories are the client's; converting ROS messages and the Lanelet2 map into the 15 input tensors is the client's; guidance services off.\n- Batch 1, one plan per request; the v5.0 capacities (320 neighbours, 140 lanes, 25 route lanes, 10 polygons, 60 line strings) and 10 DPM-Solver steps are compiled into the trace.\n- Accuracy is agreement with the fp32 CPU reference (99 scenes); no dataset-level planning metric: the paper's nuPlan closed-loop benchmark was not run, and the weights were trained on non-public TIER IV data.\n- ETH dispatch, single chip: no multi-chip mesh; WORKER dispatch (11x10) is an A/B opt-in.\n- Not a certified Autoware component; not for safety-critical driving decisions.\n", + "risks": "Device numerics (split hi / lo matmuls, fp32 LayerNorm and attention, bf16 weights elsewhere) move the plan slightly from the fp32 CPU reference: up to 0.31 m (max) / 0.14 m (mean over the 8 s) on the worst of 99 scenes, a few cm on most; the sensitivity is chaotic per scene, so any numerics change must be re-checked on all scenes. The driving quality is the weights' (trained by TIER IV on non-public data): on nuScenes, a domain they never saw, the plans are plausible but conservative. A client that skips the node's state (turn-indicator hold window, agent and ego histories, RTC prefix) gets different behaviour than the Autoware node.\n", + "licensing": "- Weights: [AutowareFoundation/diffusion_planner](https://huggingface.co/AutowareFoundation/diffusion_planner) at tag `v5.0` (commit `423efde67f5414734da43a7ad856c17ceb8b51aa`), Apache-2.0 per its model card; not redistributed here. The upstream card states that TIER IV trained the models on TIER IV synthetic and real driving data; the dataset composition is not publicly documented.\n- Pre-/post-processing ported from autoware_universe `planning/autoware_diffusion_planner` @ `9ceaccf026c31ffc5319bc9eeb4bd7bede0af3fd` (Apache-2.0).\n- Port and serving code (`code/`): Apache-2.0.\n", + "related": "Other Autoware models on Blackhole: the \"Autoware\" collection of https://huggingface.co/changh95\n", + "license": { + "id": "apache-2.0", + "name": null, + "link": null + }, + "pipeline_tag": "robotics", + "base_model": [ + "AutowareFoundation/diffusion_planner" + ] + } + } +} \ No newline at end of file