DeepLabV3Plus-R50-ONNX / int8 /benchmarks /x5h_mwmx_npu_apm80_4core.yaml
Artem Plastinkin
Add benchmarking
72c3e54
Raw History Blame
3.16 kB
hardware:
vendor: renesas
chip: rcar-x5h
cpu: arm-cortex-a720
npu: npx6-48k
npu_count: 2
npu_cores: 12
npu_default_freq_mhz: 1066
accelerator:
- npu
runtime:
engine: mwmx
toolchain_version: "MWMX SDK v4.35.0"
format: onnx
execution_provider: npu
execution_precision: int8
configuration:
npu_instances: 1
npu_cores_per_instance: 4
npu_freq_mhz: 850
benchmark:
type: hil
parameters:
batch_size: 1
input_resolution: [1, 3, 512, 1024] # explicit in the source checkpoint name (512x1024, Cityscapes)
performance:
fps: null
latency: 230.27 # synced to mwmx2.2 model_list.html (2026-09-22); no prior local measurement for this core count
# of a 4-way split model ("custom_seg_split_4_split_2") β€” not full end-to-end latency.
metrics:
accuracy: null
top5_accuracy: null
memory:
peak_mb: null
power:
avg_w: null
# Exact commands verified against the NNAC "Getting Started" chapter. Rendered
# by the AI-Dashboard in place of the generic placeholder flow β€” see
# downloadRunHTML() / parse_reproduce() in AI-Dashboard/app.js.
# CAVEAT: this artifact is segment "split_2" of a 4-way split network (see
# README) β€” these steps reproduce only this segment's latency, not an
# end-to-end DeepLabV3+ result. The network config below is required for this
# model β€” there is no working default.
reproduce:
steps:
- title: Activate the Python environment
command: >-
Activate the Python virtual environment that has the `hf` CLI
(huggingface_hub) and the NNAC toolchain installed, e.g.
`source nnac_venv/bin/activate` -- path depends on your toolchain install.
kind: note
- title: Download the ONNX model and compile config
command: hf download Renesas/DeepLabV3Plus-R50-ONNX --repo-type model --include "fp32/*" "compile_config/*" --local-dir ./DeepLabV3Plus-R50-ONNX-fp32
- title: Compile with the NNAC toolchain (INT8 auto-cast from the FP32 graph)
command: |
python3 nnac_frontend/legalize.py -d binary/nnx ./DeepLabV3Plus-R50-ONNX-fp32/fp32/deeplabv3plus_r50_oss_sim_inf.onnx --num-core 4 --network-config ./DeepLabV3Plus-R50-ONNX-fp32/compile_config/network_config.yaml
- title: Set up the R-Car X5H board
command: Configure the board per the AI Compiler (NNAC) "Getting Started" guide, section 3.4 (host TFTP/NFS setup, bootloader flashing, U-Boot, Linux boot, login) -- exact steps depend on your board/network setup.
kind: note
- title: Copy the compiled artifact to the board
command: Copy ${WORKDIR}/binary/nnx (the working directory from the download/compile steps above) to the board -- method may vary (NFS mount, scp, USB, etc.).
kind: note
- title: Run on R-Car X5H (single NPU cluster, 4 AI cores)
command: |
cd binary
./host_app ./arc_prog_npus ./nnx/deeplabv3plus_r50_oss_sim_inf
expected: hash[n] = 0x...(OK) means the run's output matches the reference hash in hash.txt; latency is the NPX execution time reported in cycles and ms.
notes: This is one segment (split_2 of 4) of the full segmentation pipeline β€” not whole-model latency.