DeepLabV3: Optimized for AMD ROCm

DeepLabV3 is a semantic segmentation model that classifies every pixel in an image into one of 21 Pascal VOC categories. This repository packages evaluation/inference for semantic segmentation using ONNX Runtime, exported and validated for AMD ROCm so it runs efficiently on AMD GPUs and CPUs.

This is based on the implementation of DeepLabV3 found here. This repository contains configurations and scripts optimized for AMD® ROCm™ platforms. You can use the DeepLabV3 AMD scripts to reproduce results or export with custom configurations. More details on model performance can be found here.


Task Overview

Task: Semantic segmentation

Dataset: Pascal VOC 2012 val (1,449 images, 21 classes including background)

Output metrics: mIoU, Pixel Accuracy, Mean Pixel Accuracy, FW-IoU, Precision, Recall, F1, Dice

NPU note: VitisAI auto-quantizes internally, so NPU runs omit an explicit precision selection.


AMD ROCm Optimization

This model export has been adapted and validated for AMD Instinct™ / Radeon™ GPUs running ROCm, as well as AMD CPUs. Key points:

  • Runs via ONNX Runtime across CPU (FP32), GPU (MIGraphX — FP32/FP16/BF16/INT8), and NPU (VitisAI — auto-quantized internally).
  • The FP32 model is exported from torchvision (DeepLabV3 MobileNetV3-Large, PASCAL VOC weights); the INT8 QDQ model is produced by vai_q_onnx static quantization for NPU.
  • No code changes required versus the upstream DeepLabV3 implementation — only environment/runtime configuration differs.
  • CPU fallback path supported for environments without a ROCm-capable GPU.
Runtime Precision Backend Hardware Notes
ONNX Runtime FP32 CPU AMD CPU Default CPU path
ONNX Runtime FP32 / FP16 / BF16 / INT8 MIGraphX AMD GPU Individually or all at once
ONNX Runtime Auto (VitisAI-quantized) VitisAI AMD NPU Auto-quantized internally

Getting Started

For setup instructions, evaluation scripts, and custom configuration options, see the DeepLabV3 on GitHub.


Model Details

Model Type: Semantic segmentation (DeepLabV3 with MobileNetV3-Large backbone)

Base Model: torchvision DeepLabV3 MobileNetV3-Large (PASCAL VOC weights)

Model Stats:

  • Input tensor: (1, 3, 513, 513), float32
  • Output tensor: (1, 21, 513, 513), float32
  • Precision tested: FP32, FP16, BF16, INT8

Accuracy Pipeline

Higher mIoU means the model's per-pixel class predictions agree more closely with ground truth across all 21 categories — 1.0 would be perfect segmentation, 0.0 means no overlap between predicted and true regions. In practice, values above ~0.6 for mIoU are considered strong for VOC-scale semantic segmentation.

Metrics Explained

Metric Description
mIoU Mean Intersection over Union — the primary segmentation metric, averaged across all 21 classes. Higher means predicted regions overlap ground truth more precisely; it penalizes both missed regions and false predictions equally, making it the most balanced single number.
Pixel Accuracy Fraction of all pixels classified correctly. Easy to saturate when background dominates — a model that predicts "background everywhere" can score high, so use mIoU as the primary metric instead.
Mean Pixel Accuracy Per-class pixel accuracy averaged across classes. Gives equal weight to rare and common classes, exposing weaknesses on small/infrequent categories that Pixel Accuracy would hide.
FW-IoU Frequency-weighted IoU — weights each class's IoU by its pixel frequency. Closer to Pixel Accuracy than mIoU; useful when you care more about getting large regions right than small ones.
Precision Mean per-class precision — of all pixels the model assigned to a class, what fraction actually belong to it. High precision means few false positives (less "hallucinated" segmentation).
Recall Mean per-class recall — of all ground-truth pixels for a class, what fraction did the model correctly label. High recall means few missed regions (false negatives).
F1 Mean per-class F1 score — harmonic mean of precision and recall. Balances the two failure modes; useful when neither false positives nor false negatives are clearly more costly.
Dice Mean per-class Dice coefficient — numerically equivalent to F1 for binary masks, included for compatibility with medical/biomedical segmentation conventions.

Accuracy Results

Full Dataset Evaluation (Pascal VOC 2012 val):

Device Precision mIoU Pixel Accuracy Mean Pixel Accuracy FW-IoU Precision Recall F1 Dice
CPU FP32 0.6748 0.9171 0.7857 0.8541 0.8111 0.7857 0.7931 0.7931
GPU FP32 0.6748 0.9171 0.7857 0.8541 0.8111 0.7857 0.7931 0.7931

Note: NPU quality may differ from CPU/GPU due to VitisAI's internal BF16 quantization.


Dig Deeper

Want to explore the full evaluation scripts, config options, and other AMD-optimized model examples?

📂 View the full project on GitHub

The GitHub repository includes:

  • Setup and prerequisites for ROCm environments
  • Benchmark, profile, and evaluation scripts for CPU, GPU, and NPU
  • Sample input/output images and annotated prediction outputs
  • Full-dataset accuracy pipeline and metrics aggregation
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support