BEVFusion (ONNX) β Renesas X5H
Introduction
This repository hosts three exported sub-graphs of BEVFusion, the multi-sensor 3D perception pipeline, targeting the Renesas R-Car X5H platform for inference on the NPX6 NPU.
These are pipeline stages, not a standalone model. Each ONNX file consumes the intermediate tensor produced by the previous stage of the full BEVFusion graph, not a raw camera or LiDAR input. The camera backbone, view transform and LiDAR branch are not included in this repository.
| Stage | File | Input (N,C,H,W) | Role |
|---|---|---|---|
| Camera neck | fp32/cam_neck.onnx |
6Γ192Γ32Γ88 |
Camera-branch FPN feature neck (6 nuScenes camera views) |
| BEV decoder | fp32/bev_decoder.onnx |
1Γ80Γ128Γ128 |
Decodes the fused/lifted BEV feature grid |
| Detection head | fp32/head.onnx |
1Γ256Γ128Γ128 |
Multi-task 3D detection head on the fused BEV backbone feature map |
- Model Architecture: BEVFusion fuses camera and LiDAR features into a shared bird's-eye-view
(BEV) representation, then runs task-specific heads on the fused BEV feature map. The head
exposes per-task outputs (
task0_heatmap,task0_reg,task0_height,task0_dim,task0_rot,task0_vel), consistent with a CenterPoint/TransFusion-style multi-task 3D detection head. The module boundaries and this interpretation are inferred from tensor shapes and output names, not independently confirmed against the official repo's code. - Paper: BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird's-Eye View Representation (Liu, Tang, Amini, Yang, Mao, Rus, Han β ICRA 2023, arXiv:2205.13542)
- Source Model: mit-han-lab/bevfusion (Apache-2.0)
- Task: 3D Object Detection (camera+LiDAR fusion, BEV), dataset: nuScenes
- Parameters: TBD β the authors publish only combined multi-task-model figures
Deployment Flow
Each FP32 ONNX stage is auto-cast to INT8 by the Renesas MWMX toolchain at compile time β no separate quantization step is required. Every stage is compiled and run separately.
cam_neck.onnx / bev_decoder.onnx / head.onnx (FP32)
β
βββΆ MWMX Runtime βββΆ INT8 auto-cast βββΆ NPX6 NPU
Provided Artifacts
| Artifact | Status | Notes |
|---|---|---|
| FP32 (ONNX) | β | fp32/cam_neck.onnx, fp32/bev_decoder.onnx, fp32/head.onnx β FP32 ONNX exports |
Performance
Measured on Renesas R-Car X5H via the MWMX runtime (APM50 ship-performance target).
Benchmark configuration: Single NPU Β· Batch size: 1 Β· Each stage run separately with its own input tensor (see the stage table above)
| AI Cores | Runtime | Precision | Device | Cam neck (ms) | BEV decoder (ms) | Head (ms) | Pipeline total (ms) | Type |
|---|---|---|---|---|---|---|---|---|
| 1 | MWMX Runtime | INT8 (auto) | X5H Β· 1Γ NPU Β· 1 Core Β· 850 MHz | 7.23 | 9.79 | 2.21 | 19.22 | Stages measured, total derived |
| 12 | MWMX Runtime | INT8 (auto) | X5H Β· 1Γ NPU Β· 12 Core Β· 850 MHz | 2.92 | 2.75 | 0.76 | 6.43 | Stages measured, total derived |
Pipeline total is the sum of the three stage latencies above, not an end-to-end run. It excludes the camera backbone, view transform and LiDAR branch, so it is not the latency of the full BEVFusion pipeline.
Accuracy
TBD β not yet measured/published for this repo.
ONNX Runtime β Renesas EP v2.1.0 (MWMX 2.2 ENG4)
Measured on R-Car X5H with ONNX Runtime + Renesas Execution Provider (INT8 auto-cast); nodes the NPU cannot run fall back to the CPU EP. Batch size 1, 1 NPU. Source: ORT Renesas EP test report, status 2026-09-30. This section supersedes earlier ORT figures in this card.
| Component | AI Cores | Latency (ms) | Throughput (fps) | NPU Inference (ms) | NPU % | CPU + Overhead % | Portable pkg NPU-only (ms) |
|---|---|---|---|---|---|---|---|
| cam_backbone | 1 | 2291.24 | 0.4 | 1911.52 | 83.4 % | 16.6 % | β |
| cam_backbone | 12 | 696.722 | 1.4 | β | β | β | β |
| vtransform_full | 1 | 2149.16 | 0.5 | 166.02 | 7.7 % | 92.3 % | β |
| vtransform_full | 12 | 2054.33 | 0.5 | β | β | β | β |
| cam_neck | 1 | 31.275 | 32 | 7.283 | 23.3 % | 76.7 % | β |
| cam_neck | 12 | 27.43 | 36.5 | β | β | β | β |
| bev_decoder | 1 | 30.331 | 33 | 9.8 | 32.3 % | 67.7 % | β |
| bev_decoder | 12 | 23.831 | 42 | β | β | β | β |
| head | 1 | 17.017 | 58.8 | 2.201 | 12.9 % | 87.1 % | β |
| head | 12 | 14.089 | 71 | β | β | β | β |
12-core latency is not always lower than 1-core: small models are dominated by CPU-side overhead.
Graph partitioning (NPU vs CPU nodes)
| Component | Total Nodes | NPU Nodes | CPU Nodes (Q/DQ inserted) | NPU % | CPU % |
|---|---|---|---|---|---|
| cam_backbone | 615 | 476 | 139 (70) | 77.4 % | 22.6 % |
| vtransform_full | 203 | 134 | 69 (4) | 66.0 % | 34.0 % |
| cam_neck | 17 | 12 | 5 (5) | 70.6 % | 29.4 % |
| bev_decoder | 44 | 42 | 2 (2) | 95.5 % | 4.5 % |
| head | 21 | 20 | 1 (1) | 95.2 % | 4.8 % |
Runtime Details
MWMX Runtime
- Engine: Renesas MWMX (Middleware MX) native inference runtime
- Input format: FP32 ONNX (compiled by the MWMX toolchain)
- NPU execution precision: INT8 (auto-cast by MWMX toolchain)
- Execution target: NPX6-48K NPU on R-Car X5H
Prerequisites
To run inference on Renesas R-Car X5H, you need:
- Renesas R-Car X5H board with NPX6 NPU
- Renesas MWMX Runtime
- Hugging Face CLI to download the models
- The remaining BEVFusion stages (camera backbone, view transform, LiDAR branch) to produce the intermediate tensors these sub-graphs expect
- A postprocessing step to decode the head's multi-task outputs (heatmap/reg/height/dim/rot/vel) into 3D bounding boxes
Download
hf download Renesas/BEVFusion-ONNX --repo-type=model --include "fp32/*"
Benchmark Methodology
- HIL runs: Hardware-in-the-loop β measured on physical R-Car X5H silicon via the MWMX runtime ("APM50" ship-performance target), each stage compiled and run separately
- Precision: FP32 ONNX input; INT8 execution (auto-cast by MWMX)
- Pipeline total: derived as the sum of the three stage latencies; no end-to-end run exists