ResNet18-OpticalFlow (ONNX) – Renesas X5H
⏳ Model file not yet uploaded. Benchmark results on this page were published ahead of the model weights — see Provided Artifacts below. Download/deployment steps will not work until the file is added to this repository.
Introduction
This repository hosts a ResNet18-based 2D optical-flow action-recognition model, targeting the Renesas R-Car X5H platform for inference on the NPX6 NPU.
- Model Architecture: ResNet18 backbone operating on stacked 2D optical-flow frames for video action recognition
- Source Model: internal export
resnet18_2d_of_hmdb5_32_{xavier,a100}(no HuggingFace mirror of these weights; seemodel.sourcein.metadata.yaml) - Task: Video Classification / Action Recognition (2D optical-flow input stream)
- Dataset: Likely HMDB51 — the
hmdb5fragment in the checkpoint names is almost certainly a truncation of "hmdb51"; confirm before publishing - Artifacts: Two artifacts are provided —
xavier-exportanda100-export. Their latencies on the X5H NPU are nearly identical (see Performance below), and the checkpoint names differ only by a reference-platform suffix (_xaviervs_a100). This strongly suggests both artifacts are the same trained checkpoint, exported/compiled with a different reference target platform (NVIDIA Jetson Xavier vs. NVIDIA A100) recorded in the source pipeline metadata — this is an observation based on the naming and near-identical performance, not a confirmed fact; the two weight files have not been diffed.
Deployment Flow
The FP32 ONNX model for each artifact is auto-cast to INT8 by the Renesas MWMX toolchain at compile time — no separate quantization step is required.
resnet18_2d_of_hmdb5_32_{xavier,a100}.onnx (FP32)
│
└─▶ MWMX Runtime ──▶ INT8 auto-cast ──▶ NPX6 NPU
Provided Artifacts
| Artifact | Parameters | Status | Notes |
|---|---|---|---|
| xavier-export (FP32 ONNX) | ~11.7M (approximate — based on the standard ResNet18 backbone; the optical-flow input adaptation may shift this slightly but not significantly) | ⏳ Not yet uploaded | Checkpoint resnet18_2d_of_hmdb5_32_xavier; benchmark numbers below exist, the model file has not been published to this repo yet |
| a100-export (FP32 ONNX) | ~11.7M (approximate — based on the standard ResNet18 backbone; the optical-flow input adaptation may shift this slightly but not significantly) | ⏳ Not yet uploaded | Checkpoint resnet18_2d_of_hmdb5_32_a100; benchmark numbers below exist, the model file has not been published to this repo yet |
Performance
Measured on Renesas R-Car X5H via the MWMX runtime (APM50 ship-performance CI pipeline).
Benchmark configuration: Single NPU · Batch size: 1 · Input resolution: not available from source data — TBD
Latencies for the two artifacts are nearly identical, consistent with them being the same underlying trained model (see the artifact note above).
| Artifact | Runtime | Precision | Device | Latency (ms) | Type |
|---|---|---|---|---|---|
| xavier-export | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | 2.163819 | Measured |
| xavier-export | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 12 Cores · 850 MHz | 1.463519 | Measured |
| a100-export | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 1 Core · 850 MHz | 2.163902 | Measured |
| a100-export | MWMX Runtime | INT8 (auto) | X5H · 1× NPU · 12 Cores · 850 MHz | 1.452456 | Measured |
Accuracy
TBD — not yet measured/published for this repo.
Runtime Details
MWMX Runtime
- Engine: Renesas MWMX (Middleware MX) native inference runtime
- Input format: FP32 ONNX (compiled by the MWMX toolchain)
- NPU execution precision: INT8 (auto-cast by MWMX toolchain)
- Execution target: NPX6-48K NPU on R-Car X5H
Prerequisites
To run inference on Renesas R-Car X5H, you need:
- Renesas R-Car X5H board with NPX6 NPU
- Renesas MWMX Runtime
- Hugging Face CLI to download the model (once the model files are published)
Download
TBD — model files not yet published to this repository.
Benchmark Methodology
- HIL runs: Hardware-in-the-loop — measured on physical R-Car X5H silicon via the MWMX
runtime (
metawaremx_runtimeCI pipeline, "APM50" ship-performance target) - Precision: FP32 ONNX input; INT8 execution (auto-cast by MWMX)
- Slices: results reported for both 1 AI core and 12 AI cores per NPU instance, for both artifacts