BEVFusion (ONNX) – Renesas X5H

Introduction

This repository hosts three exported sub-graphs of BEVFusion, the multi-sensor 3D perception pipeline, targeting the Renesas R-Car X5H platform for inference on the NPX6 NPU.

These are pipeline stages, not a standalone model. Each ONNX file consumes the intermediate tensor produced by the previous stage of the full BEVFusion graph, not a raw camera or LiDAR input. The camera backbone, view transform and LiDAR branch are not included in this repository.

Stage File Input (N,C,H,W) Role
Camera neck fp32/cam_neck.onnx 6Γ—192Γ—32Γ—88 Camera-branch FPN feature neck (6 nuScenes camera views)
BEV decoder fp32/bev_decoder.onnx 1Γ—80Γ—128Γ—128 Decodes the fused/lifted BEV feature grid
Detection head fp32/head.onnx 1Γ—256Γ—128Γ—128 Multi-task 3D detection head on the fused BEV backbone feature map
  • Model Architecture: BEVFusion fuses camera and LiDAR features into a shared bird's-eye-view (BEV) representation, then runs task-specific heads on the fused BEV feature map. The head exposes per-task outputs (task0_heatmap, task0_reg, task0_height, task0_dim, task0_rot, task0_vel), consistent with a CenterPoint/TransFusion-style multi-task 3D detection head. The module boundaries and this interpretation are inferred from tensor shapes and output names, not independently confirmed against the official repo's code.
  • Paper: BEVFusion: Multi-Task Multi-Sensor Fusion with Unified Bird's-Eye View Representation (Liu, Tang, Amini, Yang, Mao, Rus, Han β€” ICRA 2023, arXiv:2205.13542)
  • Source Model: mit-han-lab/bevfusion (Apache-2.0)
  • Task: 3D Object Detection (camera+LiDAR fusion, BEV), dataset: nuScenes
  • Parameters: TBD β€” the authors publish only combined multi-task-model figures

Deployment Flow

Each FP32 ONNX stage is auto-cast to INT8 by the Renesas MWMX toolchain at compile time β€” no separate quantization step is required. Every stage is compiled and run separately.

cam_neck.onnx / bev_decoder.onnx / head.onnx (FP32)
        β”‚
        └─▢  MWMX Runtime  ──▢  INT8 auto-cast  ──▢  NPX6 NPU

Provided Artifacts

Artifact Status Notes
FP32 (ONNX) βœ… fp32/cam_neck.onnx, fp32/bev_decoder.onnx, fp32/head.onnx β€” FP32 ONNX exports

Performance

Measured on Renesas R-Car X5H via the MWMX runtime (APM50 ship-performance target).

Benchmark configuration: Single NPU Β· Batch size: 1 Β· Each stage run separately with its own input tensor (see the stage table above)

AI Cores Runtime Precision Device Cam neck (ms) BEV decoder (ms) Head (ms) Pipeline total (ms) Type
1 MWMX Runtime INT8 (auto) X5H Β· 1Γ— NPU Β· 1 Core Β· 850 MHz 7.23 9.79 2.21 19.22 Stages measured, total derived
12 MWMX Runtime INT8 (auto) X5H Β· 1Γ— NPU Β· 12 Core Β· 850 MHz 2.92 2.75 0.76 6.43 Stages measured, total derived

Pipeline total is the sum of the three stage latencies above, not an end-to-end run. It excludes the camera backbone, view transform and LiDAR branch, so it is not the latency of the full BEVFusion pipeline.

Accuracy

TBD β€” not yet measured/published for this repo.


ONNX Runtime – Renesas EP v2.1.0 (MWMX 2.2 ENG4)

Measured on R-Car X5H with ONNX Runtime + Renesas Execution Provider (INT8 auto-cast); nodes the NPU cannot run fall back to the CPU EP. Batch size 1, 1 NPU. Source: ORT Renesas EP test report, status 2026-09-30. This section supersedes earlier ORT figures in this card.

Component AI Cores Latency (ms) Throughput (fps) NPU Inference (ms) NPU % CPU + Overhead % Portable pkg NPU-only (ms)
cam_backbone 1 2291.24 0.4 1911.52 83.4 % 16.6 % β€”
cam_backbone 12 696.722 1.4 β€” β€” β€” β€”
vtransform_full 1 2149.16 0.5 166.02 7.7 % 92.3 % β€”
vtransform_full 12 2054.33 0.5 β€” β€” β€” β€”
cam_neck 1 31.275 32 7.283 23.3 % 76.7 % β€”
cam_neck 12 27.43 36.5 β€” β€” β€” β€”
bev_decoder 1 30.331 33 9.8 32.3 % 67.7 % β€”
bev_decoder 12 23.831 42 β€” β€” β€” β€”
head 1 17.017 58.8 2.201 12.9 % 87.1 % β€”
head 12 14.089 71 β€” β€” β€” β€”

12-core latency is not always lower than 1-core: small models are dominated by CPU-side overhead.

Graph partitioning (NPU vs CPU nodes)

Component Total Nodes NPU Nodes CPU Nodes (Q/DQ inserted) NPU % CPU %
cam_backbone 615 476 139 (70) 77.4 % 22.6 %
vtransform_full 203 134 69 (4) 66.0 % 34.0 %
cam_neck 17 12 5 (5) 70.6 % 29.4 %
bev_decoder 44 42 2 (2) 95.5 % 4.5 %
head 21 20 1 (1) 95.2 % 4.8 %

Runtime Details

MWMX Runtime

  • Engine: Renesas MWMX (Middleware MX) native inference runtime
  • Input format: FP32 ONNX (compiled by the MWMX toolchain)
  • NPU execution precision: INT8 (auto-cast by MWMX toolchain)
  • Execution target: NPX6-48K NPU on R-Car X5H

Prerequisites

To run inference on Renesas R-Car X5H, you need:

  1. Renesas R-Car X5H board with NPX6 NPU
  2. Renesas MWMX Runtime
  3. Hugging Face CLI to download the models
  4. The remaining BEVFusion stages (camera backbone, view transform, LiDAR branch) to produce the intermediate tensors these sub-graphs expect
  5. A postprocessing step to decode the head's multi-task outputs (heatmap/reg/height/dim/rot/vel) into 3D bounding boxes

Download

hf download Renesas/BEVFusion-ONNX --repo-type=model --include "fp32/*"

Benchmark Methodology

  • HIL runs: Hardware-in-the-loop β€” measured on physical R-Car X5H silicon via the MWMX runtime ("APM50" ship-performance target), each stage compiled and run separately
  • Precision: FP32 ONNX input; INT8 execution (auto-cast by MWMX)
  • Pipeline total: derived as the sum of the three stage latencies; no end-to-end run exists
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Paper for Renesas/BEVFusion-ONNX