SwinV2-Tiny (ONNX) – Renesas X5H
Introduction
This repository hosts SwinV2-Tiny, targeting the Renesas R-Car X5H platform for image-classification inference on the NPX6 NPU.
- Model Architecture: Swin Transformer V2 (Tiny, window size 8, 256x256 input) — a hierarchical vision transformer using shifted local windows for self-attention, with V2's residual-post-norm and scaled cosine attention for improved training stability
- Source Model: onnx/models (timm export, swinv2_tiny_window8_256)
- Task: image-classification (dataset: imagenet-1k)
Deployment Flow
The FP32 ONNX model is auto-cast to INT8 by the Renesas MWMX toolchain at compile time — no separate quantization step is required.
swinv2_tiny_window8_256_Opset16.onnx (FP32)
│
└─▶ MWMX Runtime ──▶ INT8 auto-cast ──▶ NPX6 NPU
Provided Artifacts
| Artifact | Status | Notes |
|---|---|---|
| FP32 (ONNX) | ⏳ Pending | fp32/swinv2_tiny_window8_256_Opset16.onnx — to be added; will be auto-cast to INT8 by the MWMX toolchain at compile time (see Deployment Flow above); no separate INT8 file will be shipped |
Performance
Measured on Renesas R-Car X5H via the MWMX runtime (APM80 ship-performance CI pipeline).
Benchmark configuration: Single NPU · Batch size: 1 · Input: 3 × 256 × 256
| AI Cores | Runtime | Precision | Device | Latency (ms) | Type |
|---|---|---|---|---|---|
| 1 | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | 34.00 | Measured |
| 3 | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 3 Core · 850 MHz | 15.47 | Measured |
| 4 | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 4 Core · 850 MHz | 14.91 | Measured |
| 6 | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 6 Core · 850 MHz | 14.70 | Measured |
| 12 | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 12 Core · 850 MHz | 15.00 | Measured |
Accuracy
TBD — not yet measured/published for this repo.
ONNX Runtime – Renesas EP v2.1.0 (MWMX 2.2 ENG4)
Measured on R-Car X5H with ONNX Runtime + Renesas Execution Provider (INT8 auto-cast); nodes the NPU cannot run fall back to the CPU EP. Batch size 1, 1 NPU. Source: ORT Renesas EP test report, status 2026-09-30. This section supersedes earlier ORT figures in this card.
| Component | AI Cores | Latency (ms) | Throughput (fps) | NPU Inference (ms) | NPU % | CPU + Overhead % | Portable pkg NPU-only (ms) |
|---|---|---|---|---|---|---|---|
| Full model | 1 | 32.676 | 30.6 | 31.762 | 97.2 % | 2.8 % | 33.8467 |
| Full model | 12 | 15.484 | 64.6 | 13.619 | 88.0 % | 12.0 % | 14.9687 |
12-core latency is not always lower than 1-core: small models are dominated by CPU-side overhead.
Graph partitioning (NPU vs CPU nodes)
| Component | Total Nodes | NPU Nodes | CPU Nodes (Q/DQ inserted) | NPU % | CPU % |
|---|---|---|---|---|---|
| Full model | 515 | 514 | 1 (1) | 99.8 % | 0.2 % |
Runtime Details
MWMX Runtime
- Engine: Renesas MWMX (Middleware MX) native inference runtime
- Input format: FP32 ONNX (compiled by the MWMX toolchain)
- NPU execution precision: INT8 (auto-cast by MWMX toolchain)
- Execution target: NPX6-48K NPU on R-Car X5H
Prerequisites
To run inference on Renesas R-Car X5H, you need:
- Renesas R-Car X5H board with NPX6 NPU
- Renesas MWMX Runtime
- Hugging Face CLI to download the model
Download
hf download Renesas/SwinV2-Tiny-ONNX --repo-type=model --include "fp32/*"
Benchmark Methodology
- HIL runs: Hardware-in-the-loop — measured on physical R-Car X5H silicon via the MWMX
runtime (
metawaremx_runtimeCI pipeline, "APM80" ship-performance target) - Precision: FP32 ONNX input; INT8 execution (auto-cast by MWMX)