STM32AI MobileNetV2 0.5 ImageNet PyTorch 224 FP32 ONNX

FP32 ONNX version of STM32AI MobileNetV2 0.5 ImageNet PyTorch 224 for image classification.

Model Files

File Purpose Format
model.onnx Downloadable converted model ONNX
model.onnx Original model; no duplicate source file ONNX
graphs/netron.png ONNX graph visualization PNG
mlir/onnx.mlir ONNX-MLIR import result MLIR text

Parameter Summary

Item Value
Precision FP32
ONNX file size 7.51 MiB
Initializer tensors 106
Stored initializer elements 1,959,408
External weight files None
Original format ONNX

Stored initializer elements includes weights, biases, quantization scales, zero-points, and other constant tensors. It is not a trainable-parameter count.

Original Model Inference

The original artifact is already ONNX, so the ONNX example below applies.

The upstream artifact is already ONNX. The model.onnx file above is the source-compatible downloadable model for this variant.

Converted ONNX Inference

pip install huggingface_hub numpy onnxruntime

import numpy as np
import onnxruntime as ort
from huggingface_hub import hf_hub_download

repo_id = "ketiswp/stm32ai-MobileNetV2-0.5-ImageNet-PyTorch-224-fp32-onnx"
model_path = hf_hub_download(repo_id=repo_id, filename="model.onnx")

options = ort.SessionOptions()
options.intra_op_num_threads = 1
options.inter_op_num_threads = 1
options.execution_mode = ort.ExecutionMode.ORT_SEQUENTIAL
options.graph_optimization_level = ort.GraphOptimizationLevel.ORT_ENABLE_ALL
session = ort.InferenceSession(
    model_path,
    sess_options=options,
    providers=["CPUExecutionProvider"],
)

dtype_by_ort_type = {
    "tensor(float)": np.float32,
    "tensor(double)": np.float64,
    "tensor(float16)": np.float16,
    "tensor(int64)": np.int64,
    "tensor(int32)": np.int32,
    "tensor(int16)": np.int16,
    "tensor(int8)": np.int8,
    "tensor(uint8)": np.uint8,
    "tensor(bool)": np.bool_,
}
feeds = {}
for item in session.get_inputs():
    shape = [dim if isinstance(dim, int) and dim > 0 else 1 for dim in item.shape]
    feeds[item.name] = np.zeros(shape, dtype=dtype_by_ort_type[item.type])

outputs = session.run(None, feeds)
print([(item.name, value.shape, str(value.dtype))
       for item, value in zip(session.get_outputs(), outputs)])

Paired Model

INT8 version

Source

Project Validation

FP32/quantized comparison, conversion results, and reproduction code

This repository also includes the Netron graph, ONNX Dialect MLIR, and static MLIR dependency graph for this model variant.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including ketiswp/stm32ai-MobileNetV2-0.5-ImageNet-PyTorch-224-fp32-onnx