SwinV2-Tiny (ONNX) – Renesas X5H

Introduction

This repository hosts SwinV2-Tiny, targeting the Renesas R-Car X5H platform for image-classification inference on the NPX6 NPU.

  • Model Architecture: Swin Transformer V2 (Tiny, window size 8, 256x256 input) — a hierarchical vision transformer using shifted local windows for self-attention, with V2's residual-post-norm and scaled cosine attention for improved training stability
  • Source Model: onnx/models (timm export, swinv2_tiny_window8_256)
  • Task: image-classification (dataset: imagenet-1k)

Deployment Flow

The FP32 ONNX model is auto-cast to INT8 by the Renesas MWMX toolchain at compile time — no separate quantization step is required.

swinv2_tiny_window8_256_Opset16.onnx (FP32)
        │
        └─▶  MWMX Runtime  ──▶  INT8 auto-cast  ──▶  NPX6 NPU

Provided Artifacts

Artifact Status Notes
FP32 (ONNX) ⏳ Pending fp32/swinv2_tiny_window8_256_Opset16.onnx — to be added; will be auto-cast to INT8 by the MWMX toolchain at compile time (see Deployment Flow above); no separate INT8 file will be shipped

Performance

Measured on Renesas R-Car X5H via the MWMX runtime (APM80 ship-performance CI pipeline).

Benchmark configuration: Single NPU · Batch size: 1 · Input: 3 × 256 × 256

AI Cores Runtime Precision Device Latency (ms) Type
1 MWMX Runtime INT8 (auto) X5H · 1× NPU · 1 Core · 850 MHz 34.00 Measured
3 MWMX Runtime INT8 (auto) X5H · 1× NPU · 3 Core · 850 MHz 15.47 Measured
4 MWMX Runtime INT8 (auto) X5H · 1× NPU · 4 Core · 850 MHz 14.91 Measured
6 MWMX Runtime INT8 (auto) X5H · 1× NPU · 6 Core · 850 MHz 14.70 Measured
12 MWMX Runtime INT8 (auto) X5H · 1× NPU · 12 Core · 850 MHz 15.00 Measured

Accuracy

TBD — not yet measured/published for this repo.

ONNX Runtime – Renesas EP v2.1.0 (MWMX 2.2 ENG4)

Measured on R-Car X5H with ONNX Runtime + Renesas Execution Provider (INT8 auto-cast); nodes the NPU cannot run fall back to the CPU EP. Batch size 1, 1 NPU. Source: ORT Renesas EP test report, status 2026-09-30. This section supersedes earlier ORT figures in this card.

Component AI Cores Latency (ms) Throughput (fps) NPU Inference (ms) NPU % CPU + Overhead % Portable pkg NPU-only (ms)
Full model 1 32.676 30.6 31.762 97.2 % 2.8 % 33.8467
Full model 12 15.484 64.6 13.619 88.0 % 12.0 % 14.9687

12-core latency is not always lower than 1-core: small models are dominated by CPU-side overhead.

Graph partitioning (NPU vs CPU nodes)

Component Total Nodes NPU Nodes CPU Nodes (Q/DQ inserted) NPU % CPU %
Full model 515 514 1 (1) 99.8 % 0.2 %


Runtime Details

MWMX Runtime

  • Engine: Renesas MWMX (Middleware MX) native inference runtime
  • Input format: FP32 ONNX (compiled by the MWMX toolchain)
  • NPU execution precision: INT8 (auto-cast by MWMX toolchain)
  • Execution target: NPX6-48K NPU on R-Car X5H

Prerequisites

To run inference on Renesas R-Car X5H, you need:

  1. Renesas R-Car X5H board with NPX6 NPU
  2. Renesas MWMX Runtime
  3. Hugging Face CLI to download the model

Download

hf download Renesas/SwinV2-Tiny-ONNX --repo-type=model --include "fp32/*"

Benchmark Methodology

  • HIL runs: Hardware-in-the-loop — measured on physical R-Car X5H silicon via the MWMX runtime (metawaremx_runtime CI pipeline, "APM80" ship-performance target)
  • Precision: FP32 ONNX input; INT8 execution (auto-cast by MWMX)
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support