Google Coral DeepLabV3 MobileNetV2 1.0 Pascal VOC UINT8 ONNX

UINT8 ONNX version of Google Coral DeepLabV3 MobileNetV2 1.0 Pascal VOC for semantic segmentation.

  • Quantization: Static UINT8 quantization, QDQ format

Model Files

File Purpose Format
model.onnx Downloadable converted model ONNX
source/model.tflite Original model TensorFlow Lite
graphs/netron.png ONNX graph visualization PNG
mlir/onnx.mlir ONNX-MLIR import result MLIR text

Parameter Summary

Item Value
Precision UINT8
ONNX file size 2.21 MiB
Initializer tensors 306
Stored initializer elements 2,097,223
External weight files None
Original format TensorFlow Lite

Stored initializer elements includes weights, biases, quantization scales, zero-points, and other constant tensors. It is not a trainable-parameter count.

Original Model Inference

pip install huggingface_hub numpy ai-edge-litert

import numpy as np
from ai_edge_litert.interpreter import Interpreter
from huggingface_hub import hf_hub_download

repo_id = "ketiswp/google-coral-DeepLabV3-MobileNetV2-1.0-PascalVOC-uint8-onnx"
model_path = hf_hub_download(repo_id=repo_id, filename="source/model.tflite")
interpreter = Interpreter(model_path=model_path, num_threads=1)

for item in interpreter.get_input_details():
    signature = [int(value) for value in item.get("shape_signature", item["shape"])]
    shape = [value if value > 0 else 1 for value in signature]
    if shape != [int(value) for value in item["shape"]]:
        interpreter.resize_tensor_input(int(item["index"]), shape, strict=False)

interpreter.allocate_tensors()
for item in interpreter.get_input_details():
    value = np.zeros(tuple(int(dim) for dim in item["shape"]), dtype=item["dtype"])
    interpreter.set_tensor(int(item["index"]), value)

interpreter.invoke()
outputs = [interpreter.get_tensor(int(item["index"]))
           for item in interpreter.get_output_details()]
print([(value.shape, str(value.dtype)) for value in outputs])

Converted ONNX Inference

pip install huggingface_hub numpy onnxruntime

import numpy as np
import onnxruntime as ort
from huggingface_hub import hf_hub_download

repo_id = "ketiswp/google-coral-DeepLabV3-MobileNetV2-1.0-PascalVOC-uint8-onnx"
model_path = hf_hub_download(repo_id=repo_id, filename="model.onnx")

options = ort.SessionOptions()
options.intra_op_num_threads = 1
options.inter_op_num_threads = 1
options.execution_mode = ort.ExecutionMode.ORT_SEQUENTIAL
options.graph_optimization_level = ort.GraphOptimizationLevel.ORT_ENABLE_BASIC
session = ort.InferenceSession(
    model_path,
    sess_options=options,
    providers=["CPUExecutionProvider"],
)

dtype_by_ort_type = {
    "tensor(float)": np.float32,
    "tensor(double)": np.float64,
    "tensor(float16)": np.float16,
    "tensor(int64)": np.int64,
    "tensor(int32)": np.int32,
    "tensor(int16)": np.int16,
    "tensor(int8)": np.int8,
    "tensor(uint8)": np.uint8,
    "tensor(bool)": np.bool_,
}
feeds = {}
for item in session.get_inputs():
    shape = [dim if isinstance(dim, int) and dim > 0 else 1 for dim in item.shape]
    feeds[item.name] = np.zeros(shape, dtype=dtype_by_ort_type[item.type])

outputs = session.run(None, feeds)
print([(item.name, value.shape, str(value.dtype))
       for item, value in zip(session.get_outputs(), outputs)])

Paired Model

FP32 version

Source

Project Validation

FP32/quantized comparison, conversion results, and reproduction code

This repository also includes the Netron graph, ONNX Dialect MLIR, and static MLIR dependency graph for this model variant.

Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including ketiswp/google-coral-DeepLabV3-MobileNetV2-1.0-PascalVOC-uint8-onnx