meteor-p150 / code /PYTHON.md
changh95's picture
tt-model push meteor-p150 (container)
51defdc verified
|
Raw History Blame Contribute Delete
12.5 kB

Python API: METEOR (TIER IV, AutowareFoundation/meteor) on Blackhole

Use this API from Python code (a pipeline, a notebook, a ROS 2 node wrapper). You do not need the HTTP server: the API and the server share the decoders, the device graph and the post-processing, so the outputs and the speed are the same.

Install

Install the package on top of an environment that already has ttnn: a tt-metal python_env at 44d66500520 with patches/tt-metal-eth-dispatch.patch and patches/tt-metal-reshape-rm-sys1419.patch applied, or the tt-model container. From the root of the model repository (the directory that holds pyproject.toml, README.md and code/):

pip install -e .                        # the Python API (numpy<2, pillow, pyyaml, onnx, huggingface_hub, safetensors, opencv-python-headless)
pip install -e ".[server,test]"         # + the HTTP server and the tests

The pip project is the repository's top-level pyproject.toml; it installs the package from code/tt_meteor (there is no pyproject.toml inside code/, because the container build copies code/ over the tt-metal tree). ttnn and torch are not declared, so pip never replaces tt-metal's own build; tt-metal's python_env already has OpenCV 4.8.1, which satisfies the opencv-python-headless>=4.8,<4.12 requirement.

The package carries tt_meteor.ttaw, the shared code of the Autoware ports to Blackhole (device open, trace runner, decoders, model base class, HTTP app), vendored at the version recorded in code/tt_meteor/ttaw/VENDORED.json (0.23.2).

You want to run Extras
the Python API none
the HTTP server (tt_meteor.server.app, see SERVING.md) server
host tests (no device; the device tests are skipped): TT_VISIBLE_DEVICES=none python -m pytest -q code/tt_meteor/tests server,test
device tests: python -m pytest -q -s code/tt_meteor/tests/test_pcc_device.py code/tt_meteor/tests/test_e2e_device.py code/tt_meteor/tests/test_variants_device.py (most of them need the fp32 goldens of the development workspace and skip without them) test

Quickstart

from tt_meteor import METEOR, load_sample

with METEOR.from_pretrained(device_id=0) as model:
    out = model(**load_sample("code/tt_meteor/samples/synthetic_8cam.json"))   # 8 cameras + calibration + ego speed + stream
print(out.to_dict()["plan"]["mode"], [d["label"] for d in out.to_dicts()])

examples/quickstart.py runs the same snippet and writes the /predict JSON and a bird's-eye view (quickstart_bev.png). The shipped sample is a synthetic test frame generated by this repository (see code/tt_meteor/samples/README.md); feed your own rig's eight images and calibration for real use.

METEOR.from_pretrained(...)

METEOR.from_pretrained(
    model_id=None,          # HF repo or a local directory with the weights files; default AutowareFoundation/meteor
    *,
    revision=None,          # default for the default repo: the validated commit 01a5f6d71df (tag v1.0)
    variant=None,           # load-time variant: "default" (the only one); default $METEOR_VARIANT
    device_id=None,         # chip to open; default $TT_DEVICE_ID or 0
    device=None,            # an already-opened ttnn device (tt_meteor.device.open_device); close() does not close it
    dispatch=None,          # "eth" (p150 target, 12x10 grid) | "worker" (A/B only, 11x10) | "auto"; default $METEOR_DISPATCH or "eth"
    num_command_queues=None,  # default $METEOR_NUM_CQS or 1 (D14: a second queue gains nothing for synchronous calls)
    weights_dir=None,       # explicit local weights directory; no Hub access
    warmup_variants="default",  # trace variants to capture now; see "Warm-up"
    verbose=False,
    **compile_params,       # load-time knobs (below), e.g. input_norm="imagenet"
) -> METEOR

What it does: resolves the weights first, so a Hub problem never claims the chip (weights_dir > $METEOR_WEIGHTS_DIR > a local model_id directory > the HF snapshot at the pinned revision, restricted to meteor_v157c3Z.onnx, meteor_v157.param.yaml, LICENSE, SHA256SUMS, with an offline fallback to the cache), opens the chip (ETH dispatch, 12×10, 1 CQ; the other open parameters are DEVICE_DEFAULTS in tt_meteor/device.py: l1_small_size 32 KiB, trace_region_size 256 MiB, overridable with METEOR_*), reads the parameters of the ONNX file by consuming node, builds the graph, then compiles and captures the metal trace. If ETH dispatch cannot open (tt-metal without the patch), it warns and falls back to WORKER dispatch.

Load-time knobs (compile_params in lower case, or the environment variable METEOR_<NAME>; read once at load):

knob default meaning
input_norm onnx onnx = the released graph's /255 (ONNX Runtime parity; decision D12); imagenet = the trained ImageNet mean / std (upstream export defect, research/meteor SPEC section 10 risk 1); with imagenet an absent camera is zeroed after normalising, as in training
depth_mean_bins log depth_mean bin centres: log-spaced as exported (parity) or linear 1 + 1.25 b m (the trained bins)
max_streams 16 host temporal states kept (seg fusion, yaw tracks, mode hysteresis); the least recently used is dropped
image_precision terms3 image branch with fp32 activations, every conv as three bf16 terms (needed by the depth / seg2d argmax gates); bf16 is an A/B only: 166 ms faster per frame, but below the depth argmax gate (0.9797 < 0.99)

Warm-up

The load runs one eager frame of the whole graph (it compiles every kernel and fills the program cache) and then captures it as ONE metal trace, frame (4,678 programs). The first call is therefore as fast as the later ones. With a warm JIT cache the load takes about 45-50 s (49.7 s measured: weights 0.3 s, graph build 2.6 s, eager frame 41.0 s, capture 1.2 s). A cold kernel cache compiles for minutes (a new dispatch configuration on the build host: ETH-2CQ 142 s, WORKER 11×10 768 s). warmup_variants="none" skips the capture (the first call then captures).

Call: model(...)

Argument Type Description
images list or mapping the cameras CAM_FRONT_WIDE, CAM_FRONT_LEFT, CAM_FRONT_RIGHT, CAM_BACK_WIDE, CAM_BACK_LEFT, CAM_BACK_RIGHT, CAM_FRONT_NARROW, CAM_BACK_NARROW, any order: CameraImages, dicts {"camera", "image", "intrinsics", "T_ref_from_camera"} or {name: image}. Raw, unrectified frames of at least 768×432; larger frames are resized with OpenCV INTER_AREA semantics (bit-exact numpy port) and K is scaled per axis. The six wide / corner cameras are required; a missing or all-zero narrow camera is absent: a zero image plus its donor's K and pose (CAM_FRONT_WIDE / CAM_BACK_WIDE), the trained 7-camera configuration
calibration dict {"cameras": {name: {"intrinsics", "T_ref_from_camera"}}} (K of the image as sent; camera optical frame -> base_link: x forward, y left, z up, origin on the road) or {"preset": name} (code/tt_meteor/calib/); inline calibration of a camera wins
ego_speed float m/s, the graph's v0 (required)
stream dict optional: {"id", "reset", "timestamp_s", "T_world_from_ego"} (or "pose": [x, y, yaw]): METEOR's host temporal post-processing per stream id (BEV seg fusion needs the pose; yaw smoothing; plan-mode hysteresis). A new id or reset=True starts fresh
runtime params keyword host post-processing only (METEOR.RUNTIME_PARAMS, METEOR's C++ renderer defaults): 3D det3d_threshold 0.15, det3d_topk 64, vehicle_threshold 0.35, vru_threshold 0.15, bev_nms_iou 0.3, bev_nms_containment 0.6, stationary_logit_threshold 0.0; 2D det2d_threshold 0.30, det2d_topk 48, det2d_hide "7" (road paint); unk2d True, unk2d_threshold (= det2d), ground_z 0.0; plan mode_hysteresis 0.35, straight_margin 1.0; seg_fuse True, thin_road_edge True, yaw_smoothing True; heads False (adds the dense heads to to_dict("npz"))

tt_meteor.load_sample(path) turns a sample manifest (samples/<name>.json: image paths relative to the file, a preset, the ego speed and a stream) into these keyword arguments.

Input types

  • Images: a path, PNG / JPEG bytes, a PIL.Image, a uint8 H×W×3 RGB array, or a float array in [0, 1].
  • Transforms: 4×4 (or 3×4) matrices, {"translation", "rotation_wxyz"} (or rotation_xyzw), or Autoware {"x", "y", "z", "roll", "pitch", "yaw"}. Intrinsics: 3×3 K, 3×4 P or {"fx", "fy", "cx", "cy"}.

Output

model(...) returns a MeteorOutput (tt_meteor.Output); out.to_dict() is the /predict JSON:

key content
detections (out.to_dicts()) 3D BEV boxes sorted by score: label (VEHICLE / VRU), label_id, score, center [x, y] and size [length, width] in metres (base_link; METEOR predicts no z, height or velocity), yaw (rad, CCW from +x), stationary, future (6 × [x, y] at 0.5 .. 3 s, the agent's best mode), future_mode
trajectory, columns the selected ego path, 6 × [x, y] at 0.5 s steps
plan mode, mode_probs (softmax of the hysteresis-adjusted logits), mode_logits, paths (3 × 6 × [x, y]), steer (rad), accel (m/s²), brake_prob, dt
detections_2d per camera: label (10 classes), label_id, score, box_xyxy (pixels at 768×432)
unknown_obstacles [x, y] of 2D "obstacle" boxes placed on the ground plane (METEOR's unk2d)
traffic_light state (none / green / yellow / red), state_id, probs
lane, lane_classes the BEV lane map, uint8 [800, 500] as PNG (0.2 m cells, row 0 = +80 m ahead, column 0 = +50 m left; classes bg, road, sidewalk, crosswalk, laneline, stopline, road_edge, marking, parking), after the optional seg fusion and road-edge thinning
stationary_head_healthy, meta METEOR's health check of the stationary head; present cameras, v0, stream_id
heads (to_dict("npz") with heads=True) the dense outputs: seg2d, depth (bins), depth_mean, occupancy (class per voxel), risk (sigmoid), stationary
timing_ms preprocess, device (tables + upload + replay + read + conversion), postprocess, total

out.path is the selected path as an array; out.lane the lane map; out.boxes3d / out.boxes2d the box objects.

Lifetime and information

  • model.close() releases the trace and the device tensors and closes the chip if the model opened it; idempotent. with calls it for you; an unclosed model is closed when Python exits.
  • model.info: weights (repo, tag, revision, path), device (dispatch, grid), variant, warm variants, warm-up times, runtime parameter defaults, the load-time knobs and the trace description.
  • Calls from several threads are safe: the device calls are serialised. One model per process per chip.
  • Rig changes. The lift tables (205 MB of fp32 bin weights + the sampling grid) are built on the host per calibration (cached per rig) and written to the chip only when the calibration changes: 182 ms for the write, so a call with a new rig takes about 1.13 s instead of about 0.9 s. A fixed rig pays it once.

Speed

Warm, batch 1, ETH dispatch, 1 CQ, 12×10, AICLK 1350 MHz, p50 of 60 calls on a shared 8-core host (OPT_BASELINE.md; the device rows do not depend on the input):

stage ms
device trace, one replay (the whole network) 556.85 (back to back 556.75: 1.80 frames/s)
upload of the eight cameras (8 MB uint8) 66.5
readback (75 MB, 5 segments) + conversion to the 19 outputs 17-48 + 65
host pre-processing / camera transpose 9-11 / 22-26
host post-processing (METEOR's C++ decode rules, temporal state) 124-131
model(...) end to end 876-918 (1.09-1.14 calls/s)

This is the first release: optimization has not started (OPT_REPORT.md).

Limits

  • Batch 1 on the chip; one model per process; eight 768×432 camera slots in METEOR's order, a 400×250 lift grid at 0.4 m and an 800×500 BEV at 0.2 m (fixed shapes of the released graph); at most 3 cameras per lift cell on the rigs validated.
  • The released graph is camera-only and single-frame (its temporal memory was baked out upstream); the temporal behaviour of METEOR's runtime is host post-processing.
  • Numerics: bf16 / fp32 on the chip; outputs differ slightly from the fp32 reference (README "Demo & Performances").